Member of Technical Staff, Research Engineer (Inference)
InflectionAbout the role
What We're Building
As Inflection embarks on a new stage of growth, we are focusing on collaborating with commercial partners to adapt and fine-tune our cutting-edge models for their unique business requirements. Our accomplishments in developing, aligning, and deploying state-of-the-art models in our high EQ consumer-facing chatbot, Pi, have established a strong foundation for success. Well-funded and equipped with ample H100 resources, we have built a robust infrastructure and efficient processes to support best-in-class finetuning. By joining our team, you'll have the opportunity to contribute your expertise while being part of a dynamic organization that values innovation and collaboration.
About Inflection
Inflection is a small, interdisciplinary AI studio. We have trained several state-of-the-art language models, including Inflection 1 and Inflection 2.5, and built a personal assistant named Pi. As a studio, we are currently focused on finetuning and deploying models for specific use cases for our commercial partners.
We believe that artificial intelligence represents the beginning of an era of exponential change. Our name Inflection embraces this moment of transformation, whilst our status as a public benefit corporation provides us with the legal mandate to prioritize the well-being and happiness of our partners, users, and wider stakeholders above all else.
About the Role
Member of Technical Staff, Research Engineer (Inference)
As part of Inflection’s commitment to deploying high-performance models for enterprise applications, our inference team ensures that these models run efficiently and effectively in real-world scenarios. Research engineers in this role focus on optimizing model inference processes, reducing latency, and improving throughput without compromising model performance, ensuring robust deployment in enterprise environments.
This is a good role for you if you:
- Have experience with deploying and optimizing LLMs for inference, both in cloud and on-prem environments.
- Are adept at using tools and frameworks for model optimization and acceleration, such as ONNX, TensorRT, or TVM.
- Enjoy troubleshooting and solving complex problems related to model performance and scaling.
- Have a deep understanding of the trade-offs involved in model inference, including hardware constraints and real-time processing requirements.
- Are proficient with PyTorch and familiar with infrastructure management tools like Docker and Kubernetes for deploying inference pipelines.
Employee Pay Disclosures
At Inflection AI, we aim to attract and retain the best employees and compensate them in a way that appropriately and fairly values their individual contributions to the company. For this role, Inflection AI estimates a starting annual base salary will fall in the range of approximately $200,000 - $350,000. This estimate can vary based on the factors described above, so the actual starting annual base salary may be above or below this range.
How We Work
We value excellence and ownership. Our organizational structure focuses on individual responsibilities rather than management hierarchies. Everyone is expected to lead by doing. We are big believers in the unreasonable effectiveness of highly talented Individual Contributors who are given all the resources, space and ownership to move fast and deliver outstanding results.
Teamwork and generosity are at our core. Our culture celebrates positive challenges, asking questions, learning and actively supporting one another. This mentality of shared respect and purposeful teamwork is key to our success. We equally value all technical and non-technical contributions.
Constructive disagreement is essential. We appreciate when team members challenge assumptions, put forward new ideas, or encourage us to move faster or slower. Openness, honesty and kindness make us great.
Feedback is our ground truth. We have a tight feedback loop between the user experience and our AI creation process. Quantitative and qualitative data drives our priorities. This goes for internal culture too. Everyone has ownership and visibility into key decisions and progress.
Writing creates accountability. Whether on internal communication tools or in team memos, we are strong communicators with a special focus on the written word.
We deeply value time to reset outside of work. We encourage one another to constantly take time to recharge and always focus on maintaining a healthy work-life balance.
Engineering at Inflection
We are a ver
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s