Jobs and Careers
TO

Director of Production Engineering

Toshiba Global Commerce Solutions
United Statesfull_timeVerifiedPosted 14 Nov 2025

About the role

Toshiba Global Commerce Solutions is seeking a Director of Production Engineering to lead the reliability backbone of our global POS, cloud, and middleware platform. This strategic role owns system availability, resilience, performance, observability, and release reliability across a distributed, mission-critical commerce ecosystem. 
 
This leader will unify Site Reliability Engineering (SRE), Resilience & Performance Engineering, Observability, and AI-driven Reliability Automation into one cohesive function. As AI accelerates development velocity, verification and reliability become the core bottlenecks—making this role a cornerstone of our engineering organization. 
 
You will partner closely with Architecture, Cloud Operations, Functional Quality Engineering, and Software Development to ensure predictable reliability, smooth releases, and dramatically fewer Sev-1/Sev-2 incidents. 

Responsibilities 

System Reliability & Uptime: 

  • Define and enforce SLO/SLA frameworks, error budgets, and release criteria 
  • Lead availability, resilience, and performance strategy across all services. 
  • Own MTTR, MTBF, incident prevention, and rollback strategies at scale. 

Unified Reliability Engineering Organization: 

  • Lead teams across SRE & L3 Engineering, Resilience & Performance 
  • Engineering, Observability & Telemetry, AI Reliability Automation. 
  • Build a culture focused on prevention over firefighting. 

Architecture-Level Reliability: 

  • Collaborate with Principal Engineers and Architects to define system guardrails, resilience patterns, and failure modes. 
  • Ensure high-quality Production Readiness Reviews (PRRs) and architectural consistency. 

Resilience & Performance Engineering: 

  • Own chaos, failover, load, stress, and soak testing strategies. 
  • Validate store-mode behavior, payment workflows, edge-device dependencies, and multi-service interactions. 

Observability & Telemetry: 

  • Ensure complete, accurate signal for logs, traces, metrics, and business health. 
  • Partner with AI systems to build intelligent anomaly detection pipelines. 

AI-Driven Release Reliability: 

  • Integrate AI-based reliability scoring, resiliency prediction, automated gating, regression analysis, and incident pattern detection. 
  • Define the path toward autonomous release reliability pipelines. 

Cross-Org Leadership: 

  • Partner with Software Development, Functional Quality Engineering, Cloud Operations, Architecture, and TPM/TPO teams. 
  • Drive multi-team initiatives and ensure readiness across complex release trains. 

 

Required Experience: 

  • Bachelors Degree in Computer Science, Engineering or 10-15 years direct experience.
  • 10–15+ years in SRE, Reliability Engineering, Production Engineering, Distributed Systems, and Performance/Resilience Engineering
  • Proven ownership of uptime and system reliability in complex distributed architectures. 
  • Expertise in distributed systems, cloud platforms (AKS, Kubernetes), observability stacks (OpenTelemetry, Grafana, App Insights, Datadog), performance tuning, fault tolerance, network fundamentals, DB/service scaling, chaos testing
  • Architectural Leadership: Experience designing resilience patterns (timeouts, retries, hedging, circuit breakers). Strong partnership with architects and senior engineers. 
  • Operational Maturity: Led SRE/on-call organizations. Defined SLOs, SLIs, and error budgets at scale. Track record of driving incident prevention culture. 
  • Leadership & Communication: Builds strong engineering t

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Toshiba Global Commerce Solutions

View company profile →