Senior Machine Learning Engineer
Red HatAbout the role
About the Job
Are you passionate about shaping the future of AI by building infrastructure that ensures large language models and AI agents are safe, reliable, and aligned with human values? Red Hat's OpenShift AI team is seeking a principal ML engineer who combines deep technical expertise with a commitment to responsible AI innovation.
As a pivotal contributor to open-source projects like Open Data Hub, KServe, TrustyAI, Kubeflow, and llama-stack, you'll be at the forefront of democratizing trustworthy AI infrastructure. These critical open-source initiatives are transforming how organizations develop, deploy, and monitor machine learning models across hybrid cloud and edge environments. Your work will directly shape the next generation of MLOps platforms, making advanced AI technologies more accessible, secure, and ethically aligned.
About the Team
In today's rapidly evolving technological landscape, AI is becoming an integral part of our lives, powering everything from daily apps to complex systems in healthcare, finance, and beyond. While this is exciting, the focus on "what AI can do" has overshadowed "how it can do it safely".
Our Team’s mission is to create reliable AI systems that humans can trust. We do this by making AI safety both practical and accessible. Practical in the sense that you can implement it today, reducing complexity to facilitate adoption by developers and organizations in real-world environments, and accessible in the sense that our tools are open source and free from vendor lock-in.
What you will do
Architect and lead development of large-scale evaluation platforms for LLMs and agents, enabling automated, reproducible, and extensible assessment of accuracy, reliability, safety, and performance across diverse domains.
Define organizational standards and metrics for LLM/agent evaluation, covering hallucination detection, factuality, bias, robustness, interpretability, and alignment drift.
Build platform components and APIs that allow product teams to integrate evaluation seamlessly into training, fine-tuning, deployment, and continuous monitoring workflows.
Design automated pipelines and benchmarks for adversarial testing, red-teaming, and stress testing of LLMs and retrieval-augmented generation (RAG) systems.
Lead initiatives in multi-dimensional evaluation, including safety (toxicity, bias, harmful outputs), grounding (retrieval correctness, source attribution), and agent behaviors (tool use, planning, trustworthiness).
Collaborate with cross-functional stakeholders (safety, product, research, infrastructure) to translate abstract evaluation goals into measurable, system-level frameworks.
Advance interpretability and observability, developing tools that allow teams to understand, debug, and explain LLM behaviors in production.
Mentor engineers and establish best practices, driving adoption of evaluation-driven development across the organization.
Influence technical roadmaps and industry direction, representing the team’s evaluation-first approach in external forums and publications
What you will bring
5+ years of ML engineering experience, with 3+ years focused on large-scale evaluation of transformer-based LLMs and/or agentic systems.
Proven experience building evaluation platforms or frameworks that operate across training, deployment, and post-deployment contexts.
Deep expertise in designing and implementing LLM evaluation metrics (factuality, hallucination detection, grounding, toxicity, robustness).
Strong background in scalable platform engineering, including APIs, pipelines, and integrations used by multiple product teams.
Demonstrated ability to bridge research and engineering, operationalizing safety and alignment techniques into production evaluati
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s