Jobs and Careers
MO

Software Engineering Intern, Cloud Inference

Modular
Los Altos, United StatesinternshipVerifiedPosted 31 Jan 2026
💰 $114,000/yr($94,000/yr$114,000/yr)

About the role

About the role:


Our Cloud Inference team focuses on building a platform to serve massive foundation models with high throughput, low latency, and maximum efficiency across diverse hardwares (NVIDIA, AMD, TPUs, and more). Our goal is to make inference not only the fastest and most scalable, but also the simplest to deploy and operate.
 LOCATION: Candidates based in the United States are welcome to apply. To support growth and collaboration, all interns will work in a hybrid capacity at our Los Altos, CA office (minimum 2 days per week on-site) with relocation assistance provided for out-of-state candidates.

What you will do:


As a Software Engineering Intern on the Cloud Inference team, you’ll contribute directly to the core components of Mammoth. You’ll work alongside engineers designing large-scale distributed systems to deploy and scale foundation models with state-of-the-art performance.Depending on your interests and skills, you may work on:
  • Efficient serving – designing high-throughput, low-latency inference services, with features such as KV-aware routing and disaggregated inference.
  • KV-cache optimizations – developing distributed KV-cache manager, KV-cache offloading, and other optimizations needed to improve cache utilization.
  • Large-model inference – solving challenges in running large frontier models (e.g., DeepSeek R1) across multiple nodes.
  • Scalable deployments – extending Kubernetes APIs and building controllers to support multi-model, multi-node, and multi-cluster deployments.

What You’ll Gain:


  • A chance to help build an inference platform from the ground up, while leveraging cutting edge optimizations needed to have state-of-the-art performance.
  • Experience building a cloud inference platform from the ground up with cutting-edge optimizations.
  • Hands on experience with large-scale AI infrastructure and model serving systems.
  • Mentorship from engineers who have built AI systems at leading companies like Google, Meta and NVIDIA, and are now rebuilding the whole AI stack from the ground up.
  • The opportunity to work on real production challenges in distributed inference with immediate impact.

What you bring to the table:

 
  • Currently pursuing a Bachelor’s or Master’s degree in Computer Science, Software Engineering, Mathematics, or related field.
  • Strong programming skills in any programming language.
  • Interest in distributed systems, cloud infrastructure, or machine learning systems.
  • Curiosity, problem-solving mindset, and ability to learn quickly in a fast-moving environment.

Helpful, but not required:


  • Familiarity with Kubernetes and cloud-native technologies.
  • Strong programming skills in Go.
  • Experience building efficient, scalable distributed systems.
  • Understanding of LLMs and common serving optimizations.


What Modular brings to the table:

  • Amazing Team. We are a progressive and agile team with some of the industry’s best engineering and product leaders.
  • Competitive Compensation. We offer very strong compensation packages, including stock options. We want people to be focused on their best work and believe in tailoring compensation plans to meet the needs of our workforce. 
  • Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA.
Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and a purpose to truly change the world. The estimated base hourly range for this role is $47.00 - $57.00 USD. The hourly rate for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future.For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply as we may have openings that are lower/higher level than the ones advertised.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Modular

View company profile →