Site Reliability Engineer
EmpowerAbout the role
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.
Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.
***Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.***
The Site Reliability Engineer will help ensure the reliability, scalability, and performance of Empower’s financial services platform. This person will support production systems that serve millions of customers, improve operational excellence across assigned services, and partner with development teams to maintain high availability, strong observability, and efficient delivery practices in a regulated environment.
What you will do:
- Own operational excellence for assigned systems and services while supporting projects of varying complexity across teams
- Participate in on-call rotations, respond to incidents, troubleshoot complex system and deployment issues, and drive resolution
- Lead postmortem processes, conduct root cause analysis, and implement preventive measures
- Establish service level indicators, build proactive monitoring and alerting, and manage observability for Kubernetes environments, including EKS
- Build, maintain, and optimize infrastructure as code across multiple AWS environments
- Manage and optimize EKS clusters to support the availability, resilience, and scalability of containerized applications in production
- Collaborate with development teams to support releases and implement scalable, resilient, and maintainable services using GitOps and progressive delivery practices
- Maintain and improve CI/CD pipelines and automation tools to reduce toil and improve operational efficiency
- Lead capacity planning and right-sizing efforts to support performance optimization and system reliability
- Document critical systems, runbooks, architecture decisions, and system behaviors, and mentor entry-level SREs on operational best practices
What you will bring:
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience
- 2 to 4 years of experience in Site Reliability Engineering, DevOps, or Systems Engineering
- Experience maintaining high availability and resiliency across AWS infrastructure components, including EKS, EC2, RDS, S3, VPC, and similar services
- Production experience with Kubernetes and containerization technologies such as Docker, including deployment, troubleshooting, and optimization
- Proficiency with infrastructure as code frameworks such as Terraform and CloudFormation
- Experience with observability platforms and the technologies, systems, and networks that affect incident detection and response
- Understanding of CI/CD principles and experience with GitLab CI, Jenkins, or equivalent tools
- Knowledge of networking fundamentals, troubleshooting, and high-availability architecture patterns
- Familiarity with GitOps workflows, incident management, and on-call practices
- Strong problem-solving skills, sound judgment, and a desire to learn
What will set you apart:
- Experience in financial services or other highly regulated industries
- Familiarity with compliance frameworks such as SOC 2 or PCI DSS
- Experience with observability and APM tools such as Datadog, AppDynamics, New Relic, or similar platforms
- Strong programming skills in shell, Go, Python, or similar languages
- Experience supporting Java Spring Boot applications
- Production experience in Kubernetes, especially EKS
- Experience with service mesh technologies such as Istio or Linkerd
- AWS, Kubernetes, or other relevant certifications
- Experience with disaster recovery and business continuity planning
- Background in site reliability engineering practices and SLO/SLI methodologies
This job description is not intended to be an exhaustive list of all duties, responsibilities and qualifications of the job. The employer has the right to revise this job description at any time. You will be evaluated in part based on your performance of the responsibilities and/or tasks listed in this job description. You may be required to perform other dut
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s