Jobs and Careers
WA

Senior Site Reliability Engineer

Walmart
United Statesfull_timeVerifiedPosted 26 Feb 2026
💰 $234,000/yr($117,000/yr$234,000/yr)

About the role

What you'll do...

Position: Senior Site Reliability Engineer

Job Location: 1375 Crossman Avenue, Sunnyvale, CA 94089

Duties: Detect and document defects bugs and errors for assigned component module and conducts analysis to determine the sources under guidance. Troubleshoot performance and availability bottlenecks for assigned application under guidance. Utilize established criteria for example probability of failure frequency of failure to measure site reliability. Monitors site reliability conditions and new reliability requirements. Assists in the design and development of a reliability program plan for a specific site environment. Applies appropriate tools services or applications for reliability prediction and other site improvements. Researches and assesses various reliability models for different site environments. Assist in creation of simple modular extensible and functional design for the product solution in adherence to the requirements. Evaluate tradeoffs while designing across multiple components in a system based on the business requirements. Convert HLD to create detailed design for specific modules components of a product system. Understand nuances of designing for disaster recovery Undertake infrastructure coding automation. Assist in creation of simple modular extensible and functional design for the product solution in adherence to the requirements. Evaluate tradeoffs while designing across multiple components in a product based on the business requirements. Convert HLD to create detailed design using mock screens pseudo codes and detailed functional logic of the modules for specific modules components of a product. Understand nuances of designing for disaster recovery. Design and create MVP to clarify requirements and design and uncover risks. Independently refine the MVP design for early defects and revised customer requirements. Adhere to all relevant coding guidelines. Create and configure minimalistic Less Complex Highly Robust and high-quality code for a component module under guidance. Maintain records by documenting program development and revisions. Stay updated on the prevalent coding languages and frameworks in the industry outside the immediate scope of delivery. Identify repetitive and routine tasks in Continuous Integration Continuous Delivery CICD Testing or any other process that can be automated. Implement telemetry features as required under guidance. Apply security policy requirements to component module during code development configuration. Work with business partners to identify and document critical applications. Interprets and follows procedures in contingency plans. Explains the contingency and disaster recovery plans for assigned environment. Executes established procedures necessary to continue operations in an emergency. Participates in the design of a minimum operating environment for a computer based facility. Suggest metrics to monitor software or system performance. Monitors current performance data to ensure compliance with defined SLOs for multiple applications systems. Determines thresholds for monitoring metrics and triggers alerts based on thresholds. Supervises specific procedures to proactively check the health of applications and infrastructure including a variety of operating systems hardware and software. Makes recommendations regarding situational awareness and alerting. Make recommendations regarding instrumentation gaps and alerting logic including a variety of operating systems hardware and software. Makes recommendations regarding situational awareness and alerting. Make recommendations regarding instrumentation gaps and alerting logic.

Minimum education and experience required: Master’s degree or equivalent in computer science, computer engineering, computer information systems, software engineering, or related area and 1 year of experience in site reliability engineering, site and system administration, infrastructure management, or related area; OR Bachelor's degree or equivalent in computer science, computer engineering, computer information systems, software engineering, or related area and 3 years of experience in site reliability engineering, site and system administration, infrastructure management, or related area.

Skills required: Experience designing and implementing performance test strategies for complex web, mobile, API, and backend systems for Jira and Confluence data center instances. Experience building and maintaining automated performance test scripts using tools including JMeter, Gatling, LoadRunner, and k6. Experience performing root cause analysis of performance issues in production and test environments for Jira

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Walmart

View company profile →