Staff Site Reliability Engineer - Cloud Operations
VisaAbout the role
Company Description
Visa is a world leader in digital payments, facilitating more than 215 billion payments transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories each year. Our mission is to connect the world through the most innovative, convenient, reliable and secure payments network, enabling individuals, businesses and economies to thrive.
When you join Visa, you join a culture of purpose and belonging – where your growth is priority, your identity is embraced, and the work you do matters. We believe that economies that include everyone everywhere, uplift everyone everywhere. Your work will have a direct impact on billions of people around the world – helping unlock financial access to enable the future of money movement.
Join Visa: A Network Working for Everyone.
Job Description
CyberSource, a Visa company, is a global leader in eCommerce payment management. CyberSource was one of the world's first payment gateways, connecting online merchants to payment networks, including Visa. Today, CyberSource offers a full-service payment management platform for eCommerce merchants, combining global payment processing, fraud management and payment security systems.
CyberSource is looking for a bright, passionate and dedicated employee to join our Operations 2nd Level Applications team. In this role you will be responsible for implementing, and maintaining highly available and scalable infrastructure on AWS, GCP and our on-premise platform. These systems are truly the backbone of our business and process millions of transactions daily for some of the most prestigious companies in the world.
Essential Functions
Collaborate with cross-functional teams to implement best practices for application and infrastructure architecture on AWS and GCP.
Maintain automated deployment, monitoring, and scaling solutions to ensure the smooth operation of our cloud-based systems.
Conduct regular performance assessments and capacity planning to identify and address potential bottlenecks and ensure optimal system performance.
Implement and maintain robust monitoring, alerting, and incident response systems to proactively identify and resolve issues.
Drive continuous improvement efforts to enhance system reliability, scalability, and efficiency.
Stay up to date with the latest trends and advancements in AWS, GCP, and site reliability engineering practices, and evaluate their potential impact on our systems.
Participate in on-call rotations and lead incident response efforts to minimize downtime and resolve critical issues.
Evaluate all infrastructure changes/maintenance and determine potential for platform impact and identify mitigating steps
Act as a single point of contact, training and champion for the Enterprise platform initiatives that impact the service line
Support Customer Support on merchant escalated issues specific to applications. (e.g. increase latency)
Provide technical guidance and mentorship to team members, promoting knowledge sharing and professional development.
This is a hybrid position. Hybrid employees can alternate time between both remote and office. Employees in hybrid roles are expected to work from the office 2-3 set days a week (determined by leadership/site), with a general guidepost of being in the office 50% or more of the time based on business needs.
Qualifications
Basic Qualifications• 5+ years of relevant work experience with a Bachelor’s Degree or at least 2 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD) or 0 years of work experience with a PhD, OR 8+ years of relevant work experience.
Preferred Qualifications
• 6 or more years of work experience with a Bachelors Degree or 4 or more years of relevant experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or up to 3 years of relevant experience with a PhD
• Bachelor’s in Computer Science or related engineering field and 3 + years of experience in a highly-available Linux environment
• Experience as a Site Reliability Engineer, with a focus on both AWS and GCP.
• Strong proficiency in infrastructure as code (IaC) concepts and tools, such as Terraform or CloudFormation, for automating infrastructure deployment.
• Experience with monitoring and logging tools, such as CloudWatch, Cloud Monitoring, and ELK Stack
• In-depth knowledge of cloud services, including compute, storage, networking, databases, and security.
• Shell/Ruby/ReactJS/Python scripting knowledge
• Proficient with Swarm and Kubernetes internal architecture, networking and container micro service architectural pattern
• Deep rooted understanding of Linux
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s