Site Reliability Engineer
EmpowerAbout the role
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.
Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.
***Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.***
We are seeking a Site Reliability Engineer (SRE) to own the reliability, availability, and operational excellence of our AWS-based data platform. This role is focused on applying core SRE principles — production engineering, incident management, root cause elimination, observability, automation, and capacity planning — to large-scale data infrastructure supporting EMR, EMR Serverless, Redshift, DynamoDB, and S3.
You will treat data pipelines and analytics platforms as production systems, designing and enforcing SLAs/SLOs for uptime, performance, scalability, and data freshness. You will lead incident response, perform deep root cause analysis, implement durable fixes, and eliminate toil through automation and infrastructure-as-code.
What you will do:
Own and improve the reliability, stability, scalability, and performance of our core data platforms and services
Provide operational support for large-scale, distributed data systems, ensuring high availability and strong SLAs
Partner closely with full-stack, data, and platform engineering teams to deliver continuous improvements
Operate and support EMR and EMR Serverless (Python/Spark) workloads and data pipelines
Support and optimize Amazon Redshift and DynamoDB in high-throughput, production environments
Design, build, and evolve monitoring, alerting, and observability frameworks with a focus on symptoms, not just outages
Lead incident response, troubleshooting production issues across the full stack and coordinating with internal and external stakeholders
Perform root cause analysis (RCA) and readiness reviews; turn findings into durable fixes and automation
Create and maintain runbooks, SOPs, and operational documentation
Collaborate with engineering teams to optimize performance, reliability, and cost
Participate in an on-call rotation to respond to incidents impacting customer-facing systems
Recommend and influence the use of AWS managed services and architectural patterns
Continuously evaluate system performance, capacity, and cost to scale efficiently
What you will bring:
4–6 years of experience building or operating systems across multiple architecture domains: application, data, integration, infrastructure, and security
4+ years of hands-on AWS experience, with strong production exposure to several of the following:
Redshift, DynamoDB, EMR, EMR Serverless, EC2, S3��
Lambda, Step Functions,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s