Staff ML Infrastructure Engineer (Compute)
General MotorsAbout the role
Job Description
About the Team:
The AI Validation Platform team owns the cloud-agnostic, reliable, and cost-efficient platform that powers GM’s AV efforts. We’re proud to serve as the infrastructure platform for teams developing autonomous vehicles (L3/L4/L5). Our platform supports the simulated validation of state-of-the-art (SOTA) machine learning models, with a focus on performance, availability, concurrency, and scalability. We enable rapid innovation and development by prioritizing high-impact, ML-centric use cases.
About the Role:
We are seeking a Staff ML Infrastructure engineer to help build and scale robust Compute platforms for Simulation, data labeling, and data generation workflows. In this role, you will focus on scaling, driving efficiency, and high utilization of cutting-edge GPUs, while also leveling up the platform’s reliability. The successful candidate will have experience building and running scalable distributed systems. They will rapidly test and promote ideas, have strong problem-solving skills, and a strong bias for action.
You will play a key role in shaping the architecture, roadmap, and user experience of a robust service supporting our ML simulation and hardware-in-loop validation needs. The ideal candidate brings experience in designing distributed systems, strong problem-solving skills, and a product mindset focused on platform efficiency and reliability. This is a high-impact opportunity to influence the future of AI infrastructure at GM.
What you’ll be doing:
Collaborate with Simulation engineers, ML engineers and researchers to understand critical workflows, parse them to platform requirements, and deliver incremental value.
Own the technical roadmap, lead technical decisions on Compute architecture, caching, capacity provisioning, and auto-scaling mechanisms.
Drive the development of monitoring, observability, and metrics to ensure reliability, performance, and resource optimization.
Proactively research and integrate frameworks, hardware accelerators, and distributed computing techniques.
Lead large-scale technical initiatives across GM’s ML infrastructure.
Raise the engineering bar through technical leadership and by establishing best practices.
At a Minimum We'd Like You To Have
8+ years of industry experience, with a focus on high performance backend services.
Strong expertise in container technologies like Docker and Kubernetes.
Strong expertise in Go, or other similar coding languages.
Experience working with cloud platforms such as GCP, Azure, or AWS.
Experience in delivering cross-functional initiatives.
Strong co
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s