Jobs and Careers
WA
Senior, Software Engineer (SRE/Dev Ops)
WalmartBentonville, United Statesfull_timeVerifiedPosted 30 May 2025
💰 $180,000/yr($90,000/yr – $180,000/yr)
About the role
Position Summary...
Walmarts Transactional System provides core transactional systems to enablesegment and technology partners in creating wonderful omni experiences with speedand leverage. We are a highly motivated group of engineers, working in an agile groupto solve sophisticated and high impact problems. This role is part of Cloud PoweredCheckout team and will build the next generation multi-tenant, client agnostic, highlyscalable, omnichannel checkout solution to seamlessly enable a frictionless customercheckout experience across all sales channels globally. We process millions of orders daily through our high-performance checkout services running in Edge and Cloud.As a Site Reliability Engineer in the CPC Team, you will work with L2, Other dependent Applications, Platform team, Dev Ops and Engineering practitioners to proactively maintain mission-critical infrastructure, cloud platforms, microservices, tools, and processes that will ensure the highest levels of availability and reliability of CPC applications.About Team:Our team works closely with our US stores and e Commerce business to better serve customers by empowering team members, stores, and merchants with technological innovation. From groceries and entertainment to sporting goods and crafts, Walmart U.S. offers an extensive selection that our customers value, whether they shop online at Walmart.com, through one of our mobile apps, or in-store. Focus areas include customers, stores and employees, in-store service, merchant tools, merchant data science, and search and personalization.What you'll do:
Incident triage, Escalation and Resolution:Triage site-impacting production issues by quantifying impact, severity and urgency, analyzing systems for quick remediation, engaging the right teams for recovery [Reduce MTTE Mean Time to Engage], and focusing on immediate restoration [ Reduce MTTR Mean Time to Restore] of large-scale enterprise systems.
Alert, Monitoring, Log analysis:Detect and analyze monitoring graphs and alerts to identify systems causing production impacts with various tools like Grafana, Prometheus, MMS, Service Now, JIRA, Dynatrace, Splunk etc[Reduce MTTD Mean Time to Detect].Enhance Alerting solutions:Design and implement Java Script for the integration of alerting tool with service API endpoints with various tools like Service Now, Spotlight, Splunk, and x Matters. Requires knowledge of: Monitoring and alerting tools; Monitoring metrics and key performance indicators (for example, availability, MTBF, MTTR); SLIs and SLOs (for example, request latency, availability, error rates, saturation); Distributed tracing; Alerting logic. To demonstrate awareness of the metrics used to monitor software or system performance. Monitors current performance data to ensureadherence to defined SLOs and SLIs for simple applications/systems. Demonstrates awareness of the different types of alerts generated by the monitoring tools. Demonstrates awareness of infrastructure and application metrics.
Disaster Recovery Planning: Requires knowledge of: Disaster recovery procedures and processes; Enterprise disaster recovery systems. To work with business partners to identify and document critical applications. Interprets and follows procedures in contingency plans. Explains the contingency and disaster recovery plans for assigned environment. Executes established procedures necessary to continue operations in an emergency. Participates in the design of a minimum operating environment for a computer-based facility.
Performance and Optimization: Requires knowledge of: Unix/Linux performance optimization tuning; Java/Node JS/Tomcat/Apache tuning and optimization; Chaos tools to utilize established criteria (for example, probability of failure, frequency of failure) to measure site reliability. Monitors site reliability conditions and new reliability requirements. Assists in the design and development of a reliability program plan for a specific site environment. Applies appropriate tools, services, or applications for reliability prediction and other site improvements. Researches and assesses various reliability models for different site environments.
Work on Product Enrichment ; Content Services projects at Walmart: Develop enterprise monitoring and utilize tooling software solutions such as Grafana, Splunk etc, to improve visibility, pro-actively detect issues and restore system availabi
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s