Principal AI Scientist – Model Optimization
FlexAIAbout the role
Join FlexAI:
FlexAI is at the forefront of revolutionizing AI computing by reengineering infrastructure at the system level. Our groundbreaking architecture, combined with sophisticated software intelligence, abstraction, and an orchestration layer, allows developers to leverage a diverse array of compute, resulting in efficient, more reliable computing at a fraction of the cost. We are seeking a skilled and experienced Principal AI Scientist.
Founded by Brijesh Tripathi and Dali Kilani, who bring experience from Nvidia, Apple, Tesla, Intel, Lifen, and Zoox, FlexAI is not just building a product – we’re shaping the future of AI. Our teams are strategically distributed across Paris, Silicon Valley, and Bangalore, united by a shared mission: to deliver more compute with less complexity.
If you're passionate about shaping the future of artificial intelligence, driving innovation, and contributing to a sustainable and inclusive AI ecosystem, FlexAI is the place for you !
Position Overview:
We are looking for a Principal AI Data Scientist with deep technical expertise in model optimization to spearhead the development of high-performance AI models derived from both open-source architectures and customer-proprietary models. You’ll focused on enhancing model efficiency, latency, and throughput across diverse hardware and deployment environments.
This role is ideal for a hands-on expert who thrives at the intersection of cutting-edge AI, systems engineering, and real-world scalability.
What you’ll do:
Model Development and Optimization:
Take ownership of adapting, compressing, and optimizing open-source or proprietary models to meet specific performance goals—such as speed, accuracy, and resource efficiency—across various edge and cloud environments.
Tech Evaluation & Customization:
Evaluate and benchmark open-source LLMs, CV, and multimodal models. Customize architectures to meet customer use-case requirements including quantization, pruning, distillation, and architecture search.
Hardware-Aware Optimization:
Work closely with performance engineers to tailor models for GPUs, CPUs, and emerging accelerators (e.g., AMD, ARM, or edge chips).
Customer Collaboration:
Partner with strategic customers to understand application needs and co-develop custom model optimization workflows and pipelines.
Cross-Functional Collaboration:
Collaborate with product, engineering, MLOps, and systems teams to ensure end-to-end delivery of robust and production-grade models.
Innovation & Technical Strategy:
Stay ahead of AI optimization trends, propose new approaches, and guide architectural decisions around tooling, frameworks, and methodology.
What you’ll need to be successful:
Master’s or Ph.D. in Computer Science, Machine Learning, or related technical field.
8–10+ years in machine learning, AI R&D, or systems optimization roles.
Deep experience with model optimization techniques: quantization (INT8/FP16), pruning, distillation, NAS, etc.
Proven track record of working with LLMs, vision models, or multimodal architectures.
Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and optimization tools (e.g., ONNX, TVM, TensorRT, Hugging Face Optimum).
Strong software engineering skills in Python and/or C++.
Demonstrated ability to translate AI research into high-performance production systems.
Excellent communication skills and a customer-first mindset.
What we offer:
- A competitive salary and benefits package, tailored to recognize your dedication and contributions.
- The opportunity to collaborate with leading experts in AI and cloud computing, learning from the best and the brightest, fostering continuous g
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s