Jobs and Careers
VI

Senior Site Reliability Engineer

Visa
United Statesfull_timeVerifiedPosted 16 Apr 2026
💰 $171,800/yr($110,700/yr$171,800/yr)

About the role

Company Description

Visa is a world leader in payments and technology, with over 259 billion payments transactions flowing safely between consumers, merchants, financial institutions, and government entities in more than 200 countries and territories each year. Our mission is to connect the world through the most innovative, convenient, reliable, and secure payments network, enabling individuals, businesses, and economies to thrive while driven by a common purpose – to uplift everyone, everywhere by being the best way to pay and be paid.

Make an impact with a purpose-driven industry leader. Join us today and experience Life at Visa.

Visa’s Technology Organization is a community of problem solvers and innovators reshaping the future of commerce. We operate the world’s most sophisticated processing networks capable of handling more than 65k secure transactions a second across 80M merchants, 15k Financial Institutions, and billions of everyday people. While working with us you’ll get to work on complex distributed systems and solve massive scale problems centered on new payment flows, business and data solutions, cyber security, and B2C platforms.

Job Description

Every time someone taps, swipes, or clicks to pay-Visa infrastructure makes it happen in milliseconds, across 200+ countries. As a Senior SRE on the Product Reliability Engineering (PRE) team, you’ll own the reliability of critical production systems, drive automation that eliminates toil at scale, and help shape how we integrate AI into our engineering practices.

This isn’t a monitoring-and-tickets role. You’ll write production code, design resilient architectures, build agentic AI tools, and lead incident response for systems that process billions of transactions. You’ll have real ownership and influence on technical direction. 

WHAT YOU’LL DO

Reliability & Incident Ownership

• Own production reliability end-to-end- SLO definition, error budget tracking, proactive risk identification, and incident command during high-severity events.

• Lead root cause analysis and postmortems that drive lasting systemic improvements. Mentor I4 engineers in on-call and diagnostic best practices.

Automation & Infrastructure

• Build production-grade automation in Python (Go/Bash where appropriate) for deployment pipelines, infrastructure provisioning, and operational workflows. Develop and maintain IaC with Terraform or Ansible.

• Enhance CI/CD pipelines and observability-design monitoring, alerting, and dashboards that give teams real-time clarity on globally distributed systems.

AI-Powered Reliability

• Build GenAI-powered tools that augment incident triage, automate runbook execution, or surface predictive reliability insights. Integrate LLMs into operational workflows.

• Bring curiosity and creative thinking to identify novel AI/ML applications the team hasn’t tried yet- we want builders who see possibilities, not just problems.

Qualifications

Basic Qualifications:

  • 2+ years of relevant work experience and a Bachelor’s degree, OR 5+ years of relevant work experience.

Preferred Qualifications:

  • 3-5 years in SRE, DevOps, or Platform Engineering with a BS in CS/SE or equivalent experience.
  • Proficient in Python; working knowledge of Go, Java, or Bash.
  • Hands-on with IaC (Terraform, Ansible, or similar) and CI/CD pipelines.
  • Strong distributed systems understanding: failure modes, resilience patterns, capacity planning.
  • Proven incident management experience — on-call rotations, incident command, postmortem facilitation.
  • Experience with observability platforms (Prometheus, Grafana, Splunk, ELK, or Datadog). Linux/Unix fluency.
  • Genuine curiosity about GenAI and agentic systems — hands-on experience is a plus, willingness to learn is a must.

Nice to Have's: 

  • Cloud platforms (AWS, GCP, Azure) and container orchestration (Kubernetes).
  • Database reliability exposure: performance tuning, replication, backup/recovery, or schema change management.
  • Hands-on with AI/ML tools: LangChain, prompt engineering, or model fine-tuning.
  • SLO frameworks, error budgets, chaos engineering, or fault injection testing.
  • A GitHub profile or side project that shows us how you think and build.

Additional Information

Work Hours: Varies upon the needs of the department.

Travel Requirements: This position requires travel 5-10% of the time.

Mental/Physical Requirements: This position will be performed in an office setting.  The position will require the incumbent

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Visa

View company profile →