Software Engineer, ML Infrastructure, Level 4
Snap Inc.About the role
Snap Inc is a technology company. We believe the camera presents the greatest opportunity to improve the way people live and communicate. Snap contributes to human progress by empowering people to express themselves, live in the moment, learn about the world, and have fun together.
The Company operates Snapchat, a visual messaging app that enhances your relationships with friends, family, and the world, and Specs Inc., a wholly-owned subsidiary dedicated to making computing more human, in addition to Bitmoji, Saturn, and other digital services.
Snap Engineering teams build fun and technically sophisticated products that reach hundreds of millions of Snapchatters around the world, every day. We’re deeply committed to the well-being of everyone in our global community, which is why our values are at the root of everything we do. We move fast, with precision, and always execute with privacy at the forefront.
We’re looking for a Software Engineer to join the ML Platform Experience team, part of the core ML Platform organization. We are an AI native team which builds the agentic user experience for building, managing, and operating foundational models at Snapchat. We utilize Python as the primary language with Java/Go as supporting languages, we build agents on ADK, langfuse, and frontier LLM models, and are responsible for many other foundational technology such as model lineage, model orchestration, model data quality, and more.
What you’ll do:
Design and optimize infrastructure systems for machine learning workloads at scale and drive reliability and efficiency improvements across Snapchat’s ML Infrastructure
Build and enhance feature generation and serving pipelines that power online inferencing and offline training data generation
Develop high-performance inference systems to ensure fast and efficient AI model serving
Build infrastructure to perform scalable ML model training, evaluation, and inference in the cloud
Develop high-performance inference systems to ensure fast and efficient AI model serving
Build comprehensive data management systems for scalable data collection, labeling, processing, and evaluation
Work closely with ML engineers to deploy cutting-edge models into production
Utilize AI tools and high velocity engineering workflows to design and ship scalable services while upholding rigorous standards for code correctness, security, and production ready quality code
Knowledge, Skills & Abilities:
Strong programming skills in Python, Java
Strong problem-solving skills with a focus on system performance, scalability, and efficiency
Good understanding of distributed systems and the infrastructure components of large-scale ML
Experience with big data processing frameworks such as Spark, Flink, or Ray
Ability to collaborate and work well with others
Proven track record of operating highly-available systems at significant scale
Ability to proactively learn new concepts and apply them at work
Adaptability in learning and applying evolving AI systems and tools to remain at the forefront of engineering trends and modern development practices
Minimum Qualifications:
Bachelor’s degree in a technical field such as computer science or equivalent experience
2+ years of post-Bachelor’s software development experience; or Master’s degree in a technical field + 1+ year of post-grad software development experience; or PhD in a relevant technical field
Experience building large scale production machine learning systems, distributed systems or big data processing
Preferred Qualifications:
Masters/PhD in a technical field such as computer science or equivalent industry experience
Experience working with ML Tra
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s