Jobs and Careers
MI

Principal Software Engineer, CoreAI

Microsoft
United Statesfull_timeVerifiedPosted 14 Apr 2026
💰 $331,200/yr($142,800/yr$331,200/yr)

About the role

Overview

CoreAI is at the forefront of Microsoft’s mission to redefine how software is built and experienced. We are responsible for building the foundational platforms, services, programming models, and developer experiences that power the next generation of applications using Generative AI. Our work enables developers and enterprises to harness the full potential of AI to create intelligent, adaptive, and transformative software.   

 

The AI Core Infrastructure team, part of AI Platform team in CoreAI Organization is responsible for large-scale, highly reliable and efficient GPU management infrastructure and the inference and training platforms that power all of Microsoft’s AI workloads, such as M365 CoPilot, Github CoPilot, Microsoft CoPilot, AI Foundry’s Inference and Fine-Tuning offering of OAI and OSS models, and many more.  

 

As a Principal Engineer on the team, you’ll shape the architecture and strategy on how customers monitor, troubleshoot, and scale their AI training workloads. You’ll work across ML infrastructure, distributed systems, and observability to power large-scale pre-training, post-training, and fine-tuning on some of the world’s largest AI supercomputers. 

 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. 

 

In alignment with our Microsoft values, we are committed to cultivating an inclusive work environment for all employees to positively impact our culture every day. 



Responsibilities

As the Principal engineer on the team, your responsibilities include: 

  • Set the roadmap and drive the execution of the training infrastructure built for AI workloads at a supercomputer scale. 

  • Design, develop and ship the backend services that power the AI workloads.  

  • Deliver deep insights that empower customers to troubleshoot and optimize their large-scale AI workloads 

  • Collaborate closely with engineers, data scientists across Microsoft’s internal research teams building models to shape the infrastructure. 

  • Leverage production telemetry to influence next-generation infrastructure design, boosting efficiency, reliability, and performance 

  • Mentor and guide engineering teams, elevating technical excellence and championing a customer-focused approach to system design. 



Qualifications

Required Qualifications: 

  • Bachelor's Degree in Computer Science or related technical field and 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Python or equivalent experience. 

 

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: 
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter. 

Preferred/Additional Qualifications in one of those areas: 

  • Excellent analytical and problem-solving skills, with the ability to extract customer pain points, synthesize ambiguous requirements, and design clear, scalable solutions. 

  • Expertise with distributed observability technologies (e.g., Prometheus, OpenTelemetry, Grafana) and 2+ years of experience designing or scaling telemetry pipelines for high-throughput production systems. 

  • Advanced, hands-on experience with production ML systems, large-scale training infrastructure, NCCL, CUDA libraries and tools. 

  • 6+ years of experience building or operating distributed systems, with a strong focus on reliability, scalability, and performance. 
  • Understanding of Docker, Kubernetes, scalable architectures, and automation for production systems. 
  • Passionate and self-motivated. Strong abi

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Microsoft

View company profile →