Jobs and Careers
VI

Senior Director - Site Reliability Engineering

Visa
Foster City, United Statesfull_timeVerifiedPosted 7 Oct 2024
💰 $192,700/yr

About the role

Company Description

Visa is a world leader in payments and technology, with over 259 billion payments transactions flowing safely between consumers, merchants, financial institutions, and government entities in more than 200 countries and territories each year. Our mission is to connect the world through the most innovative, convenient, reliable, and secure payments network, enabling individuals, businesses, and economies to thrive while driven by a common purpose – to uplift everyone, everywhere by being the best way to pay and be paid.

Make an impact with a purpose-driven industry leader. Join us today and experience Life at Visa.

Job Description

Job Summary: As the Senior Director of Site Reliability Engineering (SRE), you will lead a team of SREs to ensure the highest level of performance and reliability of our services. You will be responsible for the end-to-end availability and performance of mission-critical services and building automation to prevent problem recurrence. The role requires a strategic leader who can create a vision for the SRE function and drive a culture of ‘automation first’ to improve the scalability and stability of our systems.

Essential Functions:

  • Lead and scale the SRE team, setting objectives and key results that align with the company’s strategic goals.

  • Develop and implement SRE policies, standards, and best practices for enterprise-wide systems.

  • Define standards for building reliable applications that are highly available and resilient.

  • Drive the adoption of a DevSecOps culture, fostering collaboration between development and operations teams.

  • Oversee the design and implementation of solutions for system monitoring, logging, alerting, and incident response.

  • Collaborate with product development teams to ensure reliability and scalability are considered at the design phase.

  • Manage on-call rotations, incident management processes, and post-mortem analyses to ensure continuous improvement.

  • Define Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets for all critical services.

  • Work closely with the security team to ensure compliance with industry standards and regulatory requirements.

  • Lead initiatives to improve CI/CD pipelines and automate infrastructure provisioning and deployment.

  • Provide technical leadership and mentorship to team members, encouraging professional growth and technical excellence.

This is a hybrid position. Hybrid employees can alternate time between both remote and office. Employees in hybrid roles are expected to work from the office 2-3 set days a week (determined by leadership/site), with a general guidepost of being in the office 50% or more of the time based on business needs.

Visa is not offering relocation assistance for this role.

Qualifications

Basic Qualifications:
12 or more years of work experience with a Bachelor’s Degree or at least 10 years of work experience with an Advanced degree (e.g. Masters/MBA /JD/MD), or a minimum of 5 years of work experience with a PhD

Preferred Qualifications
15 or more years of experience with a Bachelor’s Degree or 12 years of experience with an Advanced Degree (e.g. Masters, MBA, JD, or MD), PhD with 9+ years of experience in Computer Science, Engineering, or a related technical field.
Minimum of 10 years in a site reliability engineering role with at least 5 years in a leadership position managing large SRE teams.
Proficiency in system design and architecture, particularly in a cloud environment.
Expertise in automation and orchestration systems like Kubernetes, Terraform, and Ansible.
Strong coding skills in languages such as Go, Python, Ruby, or Java.
Deep understanding of networking concepts and protocols.
Experience with continuous integration and continuous deployment (CI/CD) pipelines and tools.
Proven track record of leading teams through complex system outages and scalability challenges.
Ability to mentor and grow an SRE team, fostering a culture of continuous learning and innovation.
Strong project management skills, with experience in Agile methodologies.
Excellent verbal and written communication abilities.
Proficient in creating technical documentation and system diagrams.
Experience presenting to C-level executives and stakeholders.
Demonstrated experience in incident management and post-mortem analysis.
Commitment to high availability, fault tolerance, and reliability in all aspects of work.
Knowledge of compliance and security best practices in a highly regulated industry.
Certifications in cloud technologies (AWS, GCP, Azure).
Contributions to open-source projects or public speaking at relevant tech conferences.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Visa

View company profile →