Jobs and Careers
FI

Principal Production Engineer

Finance of America
Remote, United States, United StatesRemotefull_timeVerifiedPosted 20 Feb 2026
💰 $220,000/yr($150,000/yr$220,000/yr)

About the role

Purpose of Role

Responsible for enterprise production reliability, operational resilience, and disaster recovery governance within a regulated Financial Services environment. Provides strategic and hands-on technical leadership across Incident Management, Problem Management, DevOps and IT Service Continuity Management (ITSCM), ensuring mission-critical systems remain stable, recoverable, and compliant with SOX IT General Controls (ITGC) and regulatory expectations. Also defines reliability standards, leads high-priority incident response, eliminates systemic risk through structured root cause remediation, and governs the strategy and implementation of disaster recovery capabilities aligned to business impact and financial reporting integrity. Partners closely with Engineering, Infrastructure, Development, Business, Risk, Compliance, and Internal Audit to protect customer trust, operational continuity, and the organization’s risk posture.

Key Responsibilities and Expectations

  • Serves as senior escalation authority for high-priority production incidents.
  • Leads coordinated response efforts to restore services within defined Service Level Objectives (SLOs).
  • Ensures documented impact assessments for financially significant systems.
  • Drives blameless post-incident reviews and track remediation through formal governance processes.
  • Matures enterprise incident response frameworks and escalation models.
  • Partners with Change Enablement, Risk, and Engineering teams to reduce production risk and improve service stability.
  • Owns the end-to-end Problem Management lifecycle, including root cause analysis, known error documentation, and permanent corrective actions.
  • Identifies systemic control weaknesses and drives remediation to prevent repeat incidents.
  • Establishes structured reporting on recurring incidents, MTTR trends, and control effectiveness.
  • Defines and governs enterprise Disaster Recovery (DR) strategy,  and plans.
  • Ensures alignment of Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO) with Business Impact Analysis (BIA).
  • Leads annual and periodic DR testing exercises, validation, and evidence documentation, and reports DR readiness and resilience metrics to senior leadership.
  • Coordinates with Business Continuity, Third-Party Risk, and Infrastructure teams to strengthen operational resilience.
  • Establishes enterprise production reliability standards and resilience frameworks.
  • Architects and enhances observability across applications and infrastructure (metrics, logs, traces).
  • Defines and monitors SLIs/SLOs, availability targets, and error budgets.
  • Drives automation to reduce operational toil and improve system scalability.
  • Partners in implementation and optimization of monitoring platforms (e.g., Datadog, New Relic, Elasticsearch, AWS native tools).
  • Integrates monitoring and alerting workflows with ITSM platforms (e.g., Jira Service Management) for automated ticketing and escalation.
  • Ensures Incident, Problem, Change, and DR processes support SOX ITGC design and operating effectiveness.
  • Maintains audit-ready documentation and evidence for regulatory and internal audit reviews.
  • Participates in control walkthroughs and audit engagements.
  • Identifies and remediates production control gaps impacting financial systems.
  • Establishes resilience metrics aligned to enterprise risk appetite and regulatory expectations.
  • Responds promptly and effectively to urgent business matters as they arise.
  • Performs other duties as assigned.

Reports To

  • VP, Technology Reliability and Release Engineering

Qualifications - Experience/Skills/Competencies

  • Minimum 10 years of relevant experience in Production Engineering, Disaster Recovery, DevOps, or Infrastructure Engineering.
  • Hands-on experience with the following tools and technologies, or comparable platforms: Observability: Datadog, New Relic, Elasticsearch, AWS CloudWatch;  Incident Management: JIRA Service Management, ITSM practices; CI/CD Tools: TeamCity, Octopus Deploy, Bitbucket, GitHub, Azure DevOps; Infrastructure: AWS (EC2, S3, Lambda, ECS, IAM, CloudFormation or Terraform); Backup and disaster recovery (DR) Tool: Rubrik.
  • Strong programming/scripting ability in one or more: Python, Bash, PowerShell, Go.
  • Experience building dashboards, KPIs, and reports for engineering and executive audiences.
  • Extensive knowledge of SRE frameworks, including SLOs, SLIs, MTTR, error budgets, and fault tolerance.
  • Extensive knowledge of Data Engineering principles, data lifecycle management, and data quality governance frameworks

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Finance of America

View company profile →