Principal Software Engineer - Model Inference
Red HatAbout the role
Red Hat® OpenShift® AI is a flexible, scalable artificial intelligence (AI) and machine learning (ML) platform that enables enterprises to create and deliver AI-enabled applications at scale across hybrid cloud environments. Built using open-source technologies, OpenShift AI provides trusted, operationally consistent capabilities for teams to experiment, serve models, and deliver innovative apps.
The OpenShift AI team seeks a Principal Software Engineer with Kubernetes and Model Inference Runtimes experience to join our rapidly growing engineering team. Our team focuses on making machine learning model deployment and monitoring seamless and scalable across the hybrid cloud and the edge. This is a fascinating opportunity to build and impact the next generation of hybrid cloud MLOps platforms.
What will you do:
Develop and maintain a high-quality, high-performing ML inference runtime platform for multi-modal and distributed model serving.
Contribute directly to upstream inference runtime communities such as vLLM, TGI, PyTorch, OpenVINO, and others.
Maintain CI/CD build pipelines for container images that allow faster, more secure, reliable, and frequent releases
Coordination and communication with various stakeholders
Applying a growth mindset by staying up to date with AI and ML advancements
What will you bring:
Highly experienced with programming in Python and PyTorch
Familiarity with model parallelization, quantization, and memory optimization using vLLM, TGI, and other inference libraries.
Experience with Python packaging such as PyPI libraries
Solid understanding of the fundamentals of model inferencing architectures
Experience with Jenkins, Git, shell scripting, and related technologies
Experience with the development of containerized applications in Kubernetes
Experience with Agile development methodologies
Experience with Cloud Computing using at least one of the following Cloud infrastructures AWS, GCP, Azure, or IBM Cloud
Ability to work across a large distributed hybrid engineering team
Following is considered a plus
Experience with open-source development is a plus
Development experience with C++ especially with the CUDA APIs
#LI-MD2
The salary range for this position is $142,140.00 - $234,500.00. Actual offer will be based on your qualifications.Pay Transparency
Red Hat determines compensation based on several factors including but not limited to job location, experience, applicable skills and training, external market value, and internal pay equity. Annual salary is one component of Red Hat’s compensation package. This position may also be eligible for bonus, commission, and/or equity. For positions with Remote-US locations, the actual salary range for the position may differ based on location but will be commensurate with job duties and relevant work experience.
About Red Hat
Red Hat is the world’s leading provider of enterprise open source software solutions, using a community-powered approach to deliver high-performing Linux, cloud, container, and Kubernetes technologies. Spread across 40+ countries, our associates work flexibly across work environments, from in-office, to office-flex, to fully remote, depending on the requirements of their role. Red Hatters are encouraged to bring their best ideas, no matter their title or tenure. We're a leader in open source because of our open and inclusive environment. We hire creative, passionate people ready to contribute their ideas, help solve complex problems, and make an impact.
Benefits
● Comprehensive medical, dental, and vision cov
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s