Jobs and Careers
RE
Senior MLOps Engineer - OpenShift AI - Remote, Ireland
Red HatRemote Ireland, IrelandRemotefull_timeVerifiedPosted 4 Apr 2025
About the role
<p><b><span>Job Summary</span></b></p><p><span> </span><span>Do you want to help shape the future of AI by building robust infrastructure and tools for developing trustworthy large language models and agentic workflows? We're seeking a software engineer who combines strong systems engineering skills with a passion for AI safety to develop frameworks that ensure AI systems behave reliably and align with human values. </span></p><p></p><p><span>The OpenShift AI team is looking for a Senior Software Engineer with Kubernetes and MLOps or LLMOps experience to join our rapidly growing engineering team. Our team’s focus is to make machine learning model deployment and monitoring seamless, scalable, and trustworthy across the hybrid cloud and the edge. This is a very exciting opportunity to build and impact the next generation of hybrid cloud MLOps platforms.</span></p><p><b><span> </span></b></p><p><span>In this role, you'll be contributing as a technical infrastructure expert for responsible AI features of the open source</span><a href="https://opendatahub.io/" rel="noopener noreferrer" target="_blank"><span> </span><u>Open Data Hub</u></a><span> project by actively participating in</span><a href="https://github.com/kserve" rel="noopener noreferrer" target="_blank"><span> </span><u>KServe</u></a><span>,</span><a href="https://github.com/trustyai-explainability" rel="noopener noreferrer" target="_blank"><span> </span><u>TrustyAI</u></a><span>,</span><a href="https://github.com/kubeflow/" rel="noopener noreferrer" target="_blank"><span> </span><u>Kubeflow</u></a><span>, and several other open source communities. You will work as part of an evolving development team to rapidly design, secure, build, test and release model serving, trustworthy AI, and model registry capabilities. The role is primarily an individual contributor who will be a key notable contributor to trustworthy AI and MLOps/LLMOps upstream communities and collaborate closely with the internal cross-functional development teams. </span><span> </span></p><p></p><p><b><span>Job Responsibilities</span></b></p><ul><li><p><span>Lead the architecture and implementation of MLOps/LLMOps systems within OpenShift AI, establishing best practices for scalability, reliability, and maintainability while actively contributing to relevant open source communities</span></p></li><li><p><span>Design and develop robust, production-grade features focused on AI trustworthiness, including model monitoring, bias detection, and explainability frameworks that integrate seamlessly with OpenShift AI</span></p></li><li><p><span>Drive technical decision-making around system architecture, technology selection, and implementation strategies for key MLOps components, with a focus on open source technologies like KServe and TrustyAI</span></p></li><li><p><span>Define and implement technical standards for model deployment, monitoring, and validation pipelines, while mentoring team members on MLOps best practices and engineering excellence</span></p></li><li><p><span>Collaborate with product management to translate customer requirements into technical specifications, architect solutions that address scalability and performance challenges, and provide technical leadership in customer-facing discussions</span></p></li><li><p><span>Lead code reviews, architectural reviews, and technical documentation efforts to ensure high code quality and maintainable systems across distributed engineering teams</span></p></li><li><p><span>Identify and resolve complex technical challenges in production environments, particularly around model serving, scaling, and reliability in enterprise Kubernetes deployments</span></p></li><li><p><span>Partner with cross-functional teams to establish technical roadmaps, evaluate build-vs-buy decisions, and ensure alignment between engineering capabilities and product vision</span></p></li><li><p><span>Provide technical mentorship to team members, including code review feedback, architecture guidance, and career development support while fostering a culture of engineering excellence </span></p></li></ul><p></p><p><b><span>Required Qualifications</span></b></p><ul><li><p><span>5+ years of software engineering experience, with at least 4 years focusing on ML/AI systems in production environments</span></p></li><li><p><span>Strong expertise in Python, with demonstrated experience building and deploying production ML systems</span></p></li><li><p><span>Deep understanding of Kubernetes and container orchestration, particularly in ML workload contexts</span></p></li><li><p><span>Extensive experience with MLOps tools and frameworks (e.g., KServe, Kubeflow, MLflow, or similar)</span></p></li><li><p><span>Track record of technical leadership in open source projects, including significant contributions and community engagement</span></p></li><li><p><span>Proven experience architecting and implementing large-scale distributed systems</span></p></li><li><p><span>Strong background
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s