Jobs and Careers
CO

Director of Production Engineering

CoreWeave
New York City, United Statesfull_timeVerifiedPosted 12 May 2025
💰 $275,000/yr($230,000/yr$275,000/yr)

About the role

CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services powering the next wave of AI. Our technology provides enterprises and leading AI labs with the most performant, efficient and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.

As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry.  

CoreWeave powers the creation and delivery of the intelligence that drives innovation. 

About the Role

As we continue to scale the CoreWeave Cloud Platform, ensuring reliability, performance, and operational efficiency in a live production environment is mission-critical. We’re seeking a Director of Production Engineering to lead and expand our SRE team and practices. This leader will play a key role in shaping the resilience and reliability of our platform, partnering closely with engineering and product teams to build systems and services that are secure, scalable, and highly efficient.

You will champion a culture of operational excellence, automation, and ownership, empowering our cloud platform to scale confidently and operate with agility.

What You’ll Do

  • Define and execute the SRE vision, strategy, and roadmap for a large-scale, distributed cloud infrastructure.
  • Lead and mentor a high-performing team of SREs, promoting a culture of ownership, collaboration, and continuous learning.
  • Champion automation-first practices, leveraging tools like Terraform, Kubernetes, and Infrastructure-as-Code to minimize toil and manual interventions.
  • Establish and evolve best practices in observability, monitoring, and alerting, ensuring the platform is proactive, not reactive.
  • Drive initiatives for incident management, postmortem culture, root cause analysis, and system hardening.
  • Collaborate with engineering, product, and customer support teams to build scalable, resilient, and self-healing systems.
  • Evolve our on-call strategy and processes to support a 24x7, globally distributed platform with minimal disruptions.

Who You Are

We’re looking for a thoughtful leader who blends technical depth with strategic vision, and thrives in fast-moving, high-growth environments. If you value clarity over complexity, mentorship over management, and resilience over rigidity, you’ll fit right in.

Minimum Qualifications

  • Bachelor’s degrees in Computer Science, Engineering, or related fields. 
  • 10+ years of engineering leadership roles within SRE, DevOps, or cloud infrastructure.
  • 5+ years in managing large-scale infrastructure-as-service in a geographically distributed, always-on environment.
  • Proven success leading 24x7 operations teams and delivering high-availability services at scale..
  • Deep expertise in automation, monitoring/observabilities, and incident response frameworks.
  • Familiarity with AI purpose-built cloud-native architectures, CI/CD systems, and performance tuning.

Additional Qualifications

  • Hands-on experience with Python, Go, Java, or Ruby for operational tooling and automation.
  • Strong track record of hiring, mentoring, and developing top-tier SRE talent in  high-growth companies..
  • Comfortable navigating cross-functional dynamics and influencing leadership across engineering, product, and support. 
  • Experience leading DevOps and reliability transformation projects, improving developer velocity and platform resilience.

Our compensation reflects the cost of labor across several US geographic markets. The base pay for this position ranges from $230,000-$275,000. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience.

What We Offer

The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.

In addition to a competitive salary, we offer a variety of benefits to support your needs, including:

  • Medical, dental, and vision insurance - 100% paid for by

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

CoreWeave

View company profile →