Staff Technical Program Manager, AI Infrastructure
General MotorsAbout the role
Job Description
About the Role
We are seeking a Staff Technical Program Manager (TPM) to lead AV ML Infrastructure programs for our autonomous driving platform. In this role, you will drive strategy and execution for large-scale ML infrastructure — including training pipelines, model lifecycle management, compute orchestration, and operational reliability — that power next-generation autonomy models. You will operate at the intersection of ML engineering, platform infrastructure, and operations, ensuring our ML systems are scalable, efficient, and production-ready to support end-to-end model development at scale.
---
Key Responsibilities
Program Leadership
Lead end-to-end strategic planning and execution for AI ML Infrastructure programs, delivering measurable improvements in training throughput, platform reliability, and model development velocity. Establish clear program objectives, milestones, and success metrics to drive predictable, high-quality delivery across multiple engineering and operations teams.
Cross-Functional Alignment
Collaborate with AI ML engineering, platform, validation, and product teams to define requirements, prioritize initiatives, and deliver solutions that improve AI development cycle performance and operational efficiency.
Technical Road mapping
Translate complex MLOps needs — from distributed training orchestration to compute resource management and pipeline scaling — into actionable multi-team execution plans with defined owners and measurable outcomes. Align long-term technical roadmaps with organizational goals, ensuring ML infrastructure evolves to support increasing model complexity, dataset scale, and training workloads.
Risk & Change Management
Identify technical, operational, and program risks early; develop mitigation strategies that protect training timelines, platform stability, and service reliability.
Scalability & Performance
Ensure AI ML operations processes and infrastructure are designed for long-term scalability, performance, and operational excellence — including monitoring, incident response, and capacity planning.
Metrics & Visibility
Define KPIs for ML platform performance, training system reliability, model training cycle time, and delivery velocity; maintain transparent dashboards and executive-ready reporting. Provide leadership with clear insights into progress, tradeoffs, and program health to support timely decision-making.
---
Required Qualifications:
- 10+ years of technical program management experience, including leadership of large, complex, multi-disciplinary programs.
- 5+ years working in ML Operations, ML infrastructure, AI platform engineering, or distributed compute environments.
- BS or MS in Engineering, Computer Science, or a related technical field.
- Experience supporting large-scale machine
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s