Senior Site Reliability Engineer, Cloud Infrastructure (Eng4)
ComcastAbout the role
Job Summary
As a Site Reliability Engineer (SRE), you will be part of the SRE team within the CONNECT OpTek team. Our team is responsible for development and support of multiple tools and applications used by Comcast field technicians to diagnose and troubleshoot issues within the Comcast nation-wide network.The SRE team is responsible for maintaining the existing systems, deploying new cloud environments, supporting our development teams, and implementing innovative solutions. You will work alongside software developers, testers, and project managers.
Job Description
Job Summary:
Your responsibilities will span the entire product life cycle, from requirements gathering, to development, to deployment, and operations support. You will be responsible for several things, including:
- Providing Infrastructure as Code solutions for a small cohesive group within Comcast
- Using Terraform to configure AWS Infrastructure, Kubernetes cluster provisioning and application provisioning
- Working with and supporting developers to help maintain/define best practices
- Configuring, watching, tuning and responding to monitoring events
- Supporting an on-call rotation with the SRE team
- Maintaining and improving CI/CD pipelines using Concourse and GoCD
- Supporting corporate initiatives (e.g., security hardening)
- Having a good time learning and working with the people on the team
As a senior engineer, you will provide technical leadership for your team including mentorship to less experienced engineers. You will take into consideration the needs of the business, operations teams, as well as the engineering team, when identifying prioritization of tasking for the team.
Core Responsibilities:
- Develops solutions to a wide range of difficult applications, problems or procedures
- Interprets internal/external business issues and recommends complete solutions based on best practices and proven technologies
- Works with other members of cross-functional teams, third party vendors, and company product managers and marketing teams to deliver quality products in a timely fashion that meet defined requirements
- Provides technical leadership and mentorship
- Is diligent about recording/documenting development and production support activities and tasks in our ticketing tool
- Ensures that project requests are properly accepted into the SRE engineering team, are worked in a timely and efficient manner, are of high quality, and smoothly follow the DevOps life cycle – continuous innovation, feedback, and improvement
- Deploy new systems and software and conduct appropriate testing to ensure successful deployment. Determine the necessary test coverage and plans as part of the deployment strategy
- Other duties and responsibilities as assigned
- Occasional on-call support is required
Preferred Qualifications
- Experience with Cloud Providers and configuring Infrastructure
Scripting experience with bash and python 3 - Experience troubleshooting applications and networking (Java, Angular, VPC’s Firewalls etc)
- AWS
- Kubernetes
- Terraform
- Scripting (bash/Python)
- Ansible
- Concourse/GoCD
- Docker
- Monitoring systems (Prometheus/AlertManager/Grafana)
- Git
- Experience with CM Tools (ie Terraform and Ansible)
- Experience with CI/CD Tools
- Understanding of, and experience with, distributed systems and how the pieces fit together.
- ECS/ECR
Disclaimer:
This information has been designed to indicate the general nature and level of work performed by employees in this role. It is not designed to contain or be interpreted as a comprehensive inventory of all duties, responsibilities
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s