Jobs and Careers
CO

Senior Manager - Performance and Benchmarking

CoreWeave
United Statesfull_timeVerifiedPosted 8 Aug 2025
💰 $275,000/yr($188,000/yr$275,000/yr)

About the role

CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services powering the next wave of AI. Our technology provides enterprises and leading AI labs with the most performant, efficient and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.

As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry.  

CoreWeave powers the creation and delivery of the intelligence that drives innovation. 

What You’ll Do:

We’re looking for a Senior Manager to build and lead CoreWeave’s Benchmarking & Performance team. You will own our end-to-end benchmarking program—MLPerf (Training & Inference) submissions and internal latency/throughput benchmarking—and partner closely with NVIDIA and the open-source community to push the state of the art. Your remit includes delivering service metrics (customer-facing SLO/SLA performance for live services) and public metrics (audited results, dashboards, and publications) that credibly demonstrate CoreWeave’s leadership.

You’ll collaborate across Inference, Training, Orchestration (SUNK, Kueue, Kubeflow), Networking, Storage, and Product/Marketing to set the bar for performance and transparency.

About the role:

  • Strategy & Leadership - Define the multi-year benchmarking strategy and roadmap; prioritize models/workloads (LLMs, diffusion, vision, speech) and hardware tiers. Build, lead, and mentor a high-performing team of performance engineers and data analysts. Establish governance for claims: documented methodologies, versioning, reproducibility, and audit trails.

  • Perf Ownership - Lead end-to-end MLPerf Inference and Training submissions: workload selection, cluster planning, runbooks, audits, and result publication. Coordinate optimization tracks with NVIDIA (CUDA, cuDNN, TensorRT/TensorRT-LLM, Triton, NCCL) to hit competitive results; drive upstream fixes where needed.

  • Internal Latency & Throughput Benchmarks - Design a Kubernetes-native, repeatable benchmarking service that exercises CoreWeave stacks across SUNK (Slurm on Kubernetes), Kueue, and Kubeflow pipelines.  Measure and report p50/p95/p99 latency, jitter, tokens/s, time-to-first-token, cold-start/warm-start, and cost-per-token/request across models, precisions (BF16/FP8/INT8), batch sizes, and GPU types. Maintain a corpus of representative scenarios (streaming, batch, multi-tenant) and data sets; automate comparisons across software releases and hardware generations.

  • Service & Public Metrics - Service metrics: Partner with product/eng to define SLOs for live services; publish customer-facing dashboards and weekly health reports by region/SKU/model. Public metrics: Operate a transparent performance portal (blogs, whitepapers, dashboards) with reproducible configs, container images, and methodology notes.

  • Tooling & Automation - Build CI/CD pipelines and K8s controllers/operators to schedule benchmarks at scale; integrate with observability stacks (Prometheus, Grafana, OpenTelemetry) and results warehouses. Implement supply-chain integrity for benchmark artifacts (SBOMs, Cosign signatures).
  • Cross-functional & Community - Partner with NVIDIA, key ISVs, and OSS projects (vLLM, Triton, KServe, PyTorch/DeepSpeed, ONNX Runtime) to co-develop optimizations and upstream improvements. Support Sales/SEs with authoritative numbers for RFPs and competitive evaluations; brief analysts and press with rigorous, defensible data.

Who You Are:

  • 10+ years in performance engineering for distributed systems/HPC/ML, with 3–5+ years managing or leading teams.
  • Direct experience running MLPerf submissions (Inference and/or Training) or equivalent audited benchmarks at scale.
  • Deep understanding of GPU performance (CUDA, NCCL, RDMA, NVLink/PCIe, memory bandwidth), model-server stacks (Triton, vLLM, TensorRT-LLM, TorchServe), and distributed training frameworks (PyTorch FSDP/DeepSpeed/Megatron-LM).
  • Proficient with Kubernetes and ML control pl

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

CoreWeave

View company profile →