Staff ML/AI Software Engineer – Autonomy Evaluation Platform
General MotorsAbout the role
Job Description
The Role:
As a Staff ML/AI Software Engineer within the Evaluation team under the Simulation, Evaluation, and Data organization, you will shape the design and delivery of the evaluation strategy for GM’s autonomous vehicle programs. In this role you’ll be responsible for directing software validation and model training, implementing data mining strategies and building metrics that measure autonomy performance in simulation and real-world environments. You will partner with ML engineers, systems engineers, and research scientists to build reliable, reproducible, and high-signal evaluation mechanisms that accelerate model iteration and improve safety, performance, and reliability across GM vehicles.
About the Organization:
The Evaluation team is dedicated to creating, maintaining, and evolving the evaluation ecosystem that underpins GM’s pursuit of safe, high-performing, and scalable driverless technology. The team delivers trusted metrics, automated workflows, and scalable tools that enable data-driven decision-making at every stage of AV development. Evaluation team members collaborate closely with Simulation, Motion, Perception, and Release teams, acting as system-level integrators and arbiters of end-to-end AV system quality. The organization’s remit includes development of test scenario libraries, deployment of continuous evaluation pipelines, and ownership of critical risk assessment and release gating processes. The team treats road, data mining, training, and metrics as equal use cases for our analytics framework and evaluation goals. By joining this team, you will guide the evolution of core evaluation platforms and frameworks, champion the interpretation and communication of system-level results, and play a central role in accelerating GM’s progress toward safe, validated AV deployment at scale.
What You’ll Do
Act as technical architect, defining technical vision and strategy and aligning the team with broader company objectives.
Design and implement scalable, reliable data pipelines and indexing/aggregation services to support model training and evaluation at scale, with strong guarantees for data quality, lineage, and reproducibility.
Leverage vision-language models (VLMs) and large language models (LLMs) to classify autonomy performance, mine critical scenarios, and prioritize validation efforts, integrating human-in-the-loop where appropriate.
Define and operationalize metrics and acceptance gates that quantify autonomy performance in simulation and on-road, integrated into CI/CD to guide release and merge decisions.
Build and maintain evaluation dashboards and reports that provide clear, explainable insights to engineering and leadership, including trend analysis, drift detection, and scenario coverage.
Maintain a high technical standard through architectural design, design reviews, and code reviews, setting patterns and best practices for the broader team.
<
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s