Jobs and Careers
EM

Senior Data Reliability Engineer AWS

Empower
Overland Park, United Statesfull_timeVerifiedPosted 31 Jul 2026
💰 $149,275/yr($105,700/yr$149,275/yr)

About the role

Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.

Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.

***Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.***

We are looking for a hands-on Data Reliability Engineer to own the reliability, stability, and operational excellence of our AWS-based data platform.

This role is focused on operating, troubleshooting, and improving production data systems, ensuring that data pipelines and analytics platforms are resilient, performant, and meet business-critical SLAs.

You will work closely with data and platform engineering teams to diagnose issues, resolve production incidents, and influence better design and operational practices across the data ecosystem.

What You Will Do

  • Own the reliability and stability of production data pipelines and data platform services
  • Diagnose and resolve data pipeline failures, delays, and data quality issues in production environments
  • Investigate issues across distributed data systems (e.g., Spark/EMR workloads, ingestion pipelines, warehouse performance)
  • Lead or support incident response, including triage, mitigation, and long-term resolution
  • Perform root cause analysis (RCA) and implement durable fixes to prevent recurrence
  • Define and improve data SLAs (freshness, latency, completeness) and ensure adherence
  • Design and enhance monitoring, alerting, and observability for data systems
  • Develop automation and tooling to reduce operational toil and improve system resilience
  • Contribute to disaster recovery (DR) and resiliency planning, including backup validation and recovery workflows
  • Partner with engineering teams to improve pipeline design, reliability, and operational readiness
  • Create and maintain runbooks, SOPs, and operational documentation
  • Participate in occasional off-hours support for production data systems when required

What You Will Bring

  • Minimum 5 years of experience working with production data platforms in AWS environments
  • Prior experience building data pipelines and seeing them through production, including exposure to real-world failures and operational challenges
  • Strong experience with Python and SQL in real data systems
  • Hands-on experience troubleshooting distributed data processing systems (e.g., Spark/EMR, Redshift, streaming systems)
  • Proven ability to debug and resolve production issues in data pipelines and data platforms
  • Experience with AWS data services (such as EMR, Redshift, DynamoDB, S3, or similar)
  • Experience handling production incidents and performing root cause analysis
  • Strong problem-solving mindset and ability to work through ambiguous production issues

What Will Set You Apart

  • Experience handling real-world data issues such as pipeline delays or failures
  • Experience with backfills and reprocessing
  • Experience with late-arriving or incomplete data
  • Experience improving observability and alerting specifically for data systems
  • Experience influencing or guiding data pipeline reliability and operational practices
  • Exposure to streaming/event-driven systems (Kafka, Kinesis, CDC patterns)
  • Experience with disaster recovery, backup validation, and resiliency testing
  • Strong communication during incidents with both technical and non-technical stakeholders

This job operates in a professional office environment.

This job description is not intended to be an exhaustive list of all duties, responsibilities and qualifications of the job. The employer has the right to revise this job description at any time. You will be evaluated in part based on your performance of the responsibilities and/or tasks listed in this job description. You may be required to perform other duties that are not included on this job description. The job description is not a contract for employment, and either you or the employer may terminate employment at any time, for any reason, as per terms and conditions of your employment contract.

What we offer you

We offer an a

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Empower

View company profile →