Jobs and Careers
JP
Lead Site Reliability Engineer
JPMorgan Chase & Co.United Statesfull_timeVerifiedPosted 26 Apr 2026
About the role
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
Job responsibilities
- Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team
- Leads initiatives to improve the reliability and stability of your team’s applications and platforms using data-driven analytics to improve service levels
- Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers
- Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology-related bottlenecks in your areas of expertise
- Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses
- Documents and shares knowledge within your organization via internal forums and communities of practice
Required qualifications, capabilities, and skills
- Formal training or certification with 5+ years in site reliability engineering in cloud-based environments.
- Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform
- Fluency in at least one programming language such as (e.g., Python, Java Spring Boot, .Net, etc.)
- Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
- Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc.
- Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.)
- Experience with container and container orchestration (e.g., ECS, Kubernetes, Docker, etc.)
- Experience with troubleshooting common networking technologies and issues
Ability to identify and solve problems related to complex data structures and algorithms
- Direct experience with SRE, incident/problem management, RCA methods and techniques, service health metrics (e.g., MTTD, MTTR), and post-incident reviews.
- Applied use of LLMs/agents, RAG, anomaly detection, or automated runbooks to accelerate evidence collection, summarization, and action routing in review workflows.
- Familiarity with structured methods used in high-reliability investigations (e.g., Bowtie/AcciMap/STPA), peer review/checklists, cross-source corroboration, cognitive bias mitigation (e.g., confirmation, hindsight, outcome bias), and evidence-handling practices such as immutable log retention, event timestamping, query capture, and “docket”-style evidence packages suitable for leadership reviews and audits.
- Experience with modern cloud data platforms and workflow orchestration (e.g., warehouses/lakehouses, streaming, Airflow/Prefect/dbt) and integration with systems like ServiceNow or Jira.
- Background in financial services or other regulated, large-scale operating models; comfort with data privacy, retention, and access controls.
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may re
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s