Senior Software Engineer, Fleet Monitoring Analysis
CoreWeaveAbout the role
What You’ll Do:
The Fleet Monitoring & Analysis (FMA) team within Fleet Engineering builds and operates the metrics, monitoring, and alerting systems that power CoreWeave’s automated provisioning and lifecycle management of our global hardware fleet. As the team behind many of the exporters, dashboards, Slack bots, and alerting rules used across Fleet Engineering and Operations, FMA plays a central role in enabling zero-touch, high-reliability operations for GPU servers and the environments they run in.
About the Role:
As an Engineer on the Fleet Monitoring & Analysis team, you’ll help build, run, and refine the metrics, alerts, visualizations, and data-driven insights that keep CoreWeave’s ever-expanding fleet of hardware nodes healthy and observable.
You’ll join a mixed-skill engineering team focused on elevating the art of managing high-performance hardware at scale—partnering closely with Fleet Engineering and Operations, and our observability platform teams to turn telemetry into automation, operational efficiency, and a better experience for CoreWeave’s customers.
In this role, you will:
- Design and implement large-scale server observability solutions that improve the stability and reliability of CoreWeave’s global hardware fleet.
- Adapt, extend, and implement open-source monitoring and alerting tooling (e.g., Prometheus-compatible exporters, AlertManager/Victoria Metrics) to deepen our visibility into fleet and environmental health.
- Generate and maintain tailored reports, alarms, and visualizations used by Fleet Engineering and FROps to understand, respond to, and plan for fleet growth and change.
- Create and evolve test plans, deployment automation, dashboards, alerts, and insights around fleet operations, and participate in the Fleet Engineering Developers’ on-call rotation.
- Collaborate with teammates across FMA to invest in each other’s growth, share ideas, and continuously improve how we monitor and automate CoreWeave’s infrastructure.
Who You Are:
- 2+ years of experience in a software or infrastructure engineering role in industry.
- Experience with automation and orchestration workflows, and familiarity with server hardware, components, and strategies for managing physical infrastructure at scale.
- Experience implementing metrics collection and alerting on standard monitoring platforms (for example, Prometheus/Victoria Metrics, AlertManager, Grafana, or similar observability stacks).
- Proficiency working in Linux-based environments and using at least one scripting or programming language (such as Python, Go, or similar) to build tooling and automation.
Preferred:
- Experience designing or operating time-series monitoring at scale using Prometheus-compatible systems, Victoria Metrics, and related observability tooling (e.g., Grafana, AlertManager).
- Experience deploying and maintaining application services in Kubernetes.
- Experience building or maintaining Slack bots, webhooks, or automation that consume alerts and drive lifecycle actions across a fleet (for example, via Kubernetes controllers, Ansible, or similar tools).
- Experience with data warehousing, SQL, and building reporting pipelines or dashboards for operational analytics in environments like Grafana or similar BI tools.
Wondering if you’re a good fit?
We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk.
- You enjoy digging into metrics, logs, and alerts to understand how large-scale systems behave over time.
- You’re excited to collaborate with operations and engineering teams to turn manual runbooks into automation and improve fleet reliability.
- You like working at the intersection of hardware, sof
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s