Senior Manager, Capacity and Performance Engineering
General MotorsAbout the role
Job Description
ABOUT THE ROLE
GM is building a next-generation autonomous vehicle system targeting highway launch. That system runs on GPU clusters for model training, on large-scale simulation infrastructure for validation, and on data pipelines that ingest and process sensor data at petabyte scale. All of it cloud-hosted, all of it needs to scale predictably as the program accelerates.
As Senior Manager of AV Cloud Capacity & Performance Engineering, you own the team and function responsible for making sure that infrastructure is always scaled, efficiently utilized, and fiscally sound. You'll build and operate the capacity modeling, cost optimization, performance benchmarking, and vendor management capabilities that allow the AV engineering organization to develop without resource ceilings.
In practice, this means making calls like: committing to a GPU procurement months before demand materializes based on program signals, calling a cross-team efficiency initiative when utilization data reveals a pattern of waste, or recommending to VP leadership that a next-generation hardware cluster is worth the migration cost based on your team's benchmark data.
WHAT YOU'LL DO
Own the compute capacity forecast the AV program depends on
Build and maintain detailed capacity models for the full AV compute stack: model training clusters, simulation infrastructure, data ingest systems, inference serving, and developer tooling. Translate upcoming program milestones into specific capacity commitments well ahead of when supply is needed. Own the GPU supply plan end-to-end: assessing when to transition workloads across hardware generations, aligning procurement timing with vendor product cycles, and ensuring supply commitments are in place before engineering demand materializes.
Run an infrastructure efficiency and cost optimization program with measurable targets
Lead a cross-team program to reduce compute and storage waste without degrading engineering velocity. Your team surfaces over-provisioned workloads, identifies GPU scheduling and utilization inefficiencies, and develops architectural recommendations that improve cost per experiment. You define the targets, own the results reporting to VP leadership, and drive execution in partnership with the engineering teams that own the workloads.
Operate the performance lab and deliver actionable benchmarks to engineering leadership
Lead the function that characterizes how AV workloads perform on the organization's infrastructure training throughput on current and next-generation GPU clusters, simulation efficiency per node-hour, and storage I/O characteristics. Establish repeatable benchmarks that answer concrete procurement and architecture questions: Is the next GPU generation worth the migration cost? Where are the actual throughput bottlenecks for training workloads? Deliver findings as engineering-grade recommendations, not observational reports.
Own strategic vendor and cloud provider relationships
Own the working relationships with major cloud providers and hardware vendors at a program level. Define and enforce SLOs for vendor responsiveness and supply commitments. Build the monitoring and escalation infrastructure to catch vendor performance issues before they become engineering blockers. Negotiate supply and pricing terms that serve a multi-year capacity roadmap.
Advise VP and Director-level leaders on infrastructure tradeoffs and program risk
Be the person senior engineering and program leaders turn to
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s