Director of Production Engineering
Toshiba Global Commerce SolutionsAbout the role
Toshiba Global Commerce Solutions is seeking a Director of Production Engineering to lead the reliability backbone of our global POS, cloud, and middleware platform. This strategic role owns system availability, resilience, performance, observability, and release reliability across a distributed, mission-critical commerce ecosystem.
This leader will unify Site Reliability Engineering (SRE), Resilience & Performance Engineering, Observability, and AI-driven Reliability Automation into one cohesive function. As AI accelerates development velocity, verification and reliability become the core bottlenecks—making this role a cornerstone of our engineering organization.
You will partner closely with Architecture, Cloud Operations, Functional Quality Engineering, and Software Development to ensure predictable reliability, smooth releases, and dramatically fewer Sev-1/Sev-2 incidents.
Responsibilities
System Reliability & Uptime:
- Define and enforce SLO/SLA frameworks, error budgets, and release criteria
- Lead availability, resilience, and performance strategy across all services.
- Own MTTR, MTBF, incident prevention, and rollback strategies at scale.
Unified Reliability Engineering Organization:
- Lead teams across SRE & L3 Engineering, Resilience & Performance
- Engineering, Observability & Telemetry, AI Reliability Automation.
- Build a culture focused on prevention over firefighting.
Architecture-Level Reliability:
- Collaborate with Principal Engineers and Architects to define system guardrails, resilience patterns, and failure modes.
- Ensure high-quality Production Readiness Reviews (PRRs) and architectural consistency.
Resilience & Performance Engineering:
- Own chaos, failover, load, stress, and soak testing strategies.
- Validate store-mode behavior, payment workflows, edge-device dependencies, and multi-service interactions.
Observability & Telemetry:
- Ensure complete, accurate signal for logs, traces, metrics, and business health.
- Partner with AI systems to build intelligent anomaly detection pipelines.
AI-Driven Release Reliability:
- Integrate AI-based reliability scoring, resiliency prediction, automated gating, regression analysis, and incident pattern detection.
- Define the path toward autonomous release reliability pipelines.
Cross-Org Leadership:
- Partner with Software Development, Functional Quality Engineering, Cloud Operations, Architecture, and TPM/TPO teams.
- Drive multi-team initiatives and ensure readiness across complex release trains.
Required Experience:
- Bachelors Degree in Computer Science, Engineering or 10-15 years direct experience.
- 10–15+ years in SRE, Reliability Engineering, Production Engineering, Distributed Systems, and Performance/Resilience Engineering
- Proven ownership of uptime and system reliability in complex distributed architectures.
- Expertise in distributed systems, cloud platforms (AKS, Kubernetes), observability stacks (OpenTelemetry, Grafana, App Insights, Datadog), performance tuning, fault tolerance, network fundamentals, DB/service scaling, chaos testing
- Architectural Leadership: Experience designing resilience patterns (timeouts, retries, hedging, circuit breakers). Strong partnership with architects and senior engineers.
- Operational Maturity: Led SRE/on-call organizations. Defined SLOs, SLIs, and error budgets at scale. Track record of driving incident prevention culture.
- Leadership & Communication: Builds strong engineering t
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s