Jobs and Careers
MA

Senior Site Reliability Engineer

Mastercard
O'Fallon, United Statesfull_timeVerifiedPosted 11 Mar 2024

About the role

Our Purpose

We work to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation and delivers better business results.

Title and Summary

Senior Site Reliability Engineer

About the Role
The Business Operations (Biz Ops) team is seeking a Senior Site Reliability Engineer (SRE).
The role of Business Operations Organization is to be the production readiness steward for Mastercard products. As a Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to run our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principals that includes operational design, automation, capacity planning, monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.

We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments.
Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle.

All About the Program You Support (Data Platform & Engineering Services):
Our Mission:
The MasterCard Enterprise Data Warehouse provides robust, high performance, and secure repositories of data assets.
Everyday, everywhere these unassailable data assets empower stakeholders to make intelligent, data driven decisions, allowing them to grow, diversify, and build business.
Our Vision:
We will become the premier data warehouse in the world by enabling cutting edge technology, analytics, business intelligence, visualization, and data science, driving value at MasterCard.

What you’ll do:
• Plan, manage, and oversee all aspects of a Production Environment for Data Platforms, Business Intelligence Platforms & Data Cloud.
• Manage, lead and coach the Site Reliability Engineers.
• Define strategies for Application Performance Monitoring, Unit Cost and Chaos Engineering aspects.
• Find ways for Continuous Optimizations in a Production Environment.
• Ability to understand MTTR, SLO, SLI definitions and apply them to services.
• Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
• Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data, services and infrastructures.
• Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
• Practice sustainable incident response and blameless postmortems.
• Ensures that batch production scheduling and process are accurate and timely.
• Able to create and execute queries to big data platform and relational data tables to identify process issues or to perform mass updates, preferred.
• Ability to isolate problems between hardware and software. Working with appropriate team(s) and vendors until a resolution has been reached.
• Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
• Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
• Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
• Maintain services on

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Mastercard

View company profile →