Senior Production Engineer, Tooling & Frameworks
CoreWeaveAbout the role
About the Role
Production Engineering ensures CoreWeave’s cloud runs with world-class reliability, performance, and operational excellence. Herd is our newest innovation: an agentic AI platform that serves as CoreWeave’s intelligent SRE assistant - combining AI reasoning, data infrastructure, and observability into an autonomous operational intelligence layer for internal use.
As a Production Engineer on Herd, you’ll define and build the systems that power a scalable agentic ecosystem. You’ll design distributed services and data pipelines that process, embed, and retrieve operational knowledge at scale, enabling LLM-powered agents to work alongside human engineers in production. This is a hands-on role at the intersection of AI operations, distributed systems, and data infrastructure.
What You’ll Do
- Architect and build large-scale distributed systems that power AI SRE Platforms.
- Design data infrastructure for AI reasoning (embedding generation, context retrieval, vector stores) optimized for real-time operational queries.
- Build agent orchestration and lifecycle components so agents can communicate, delegate, and reason collectively across CoreWeave systems.
- Integrate AI SRE Platform with a large number of internal systems (Kubernetes, observability platforms, etc.) to enable end-to-end automation and insights.
- Lead architectural design discussions and set technical direction for AI-driven reliability systems.
- Partner across Production Engineering, Data Engineering, ML Infrastructure, and Platform to operate AI SRE as a high-availability platform embedded in critical reliability workflows.
- Develop services that interpret telemetry, detect anomalies, and generate RCA (root cause analysis) and PIR (post-incident review) artifacts; trigger automated mitigations where appropriate.
- Codify operational best practices into services, APIs, and Kubernetes-native components.
- Participate in an on-call rotation supporting the systems you build.
What You’ve Worked On (Minimum Qualifications)
- 5+ years in software or infrastructure engineering building and operating distributed systems at scale.
- Proficiency in Python (or similar), with experience delivering production microservices and data pipelines.
- Expertise in Kubernetes, container orchestration, and cloud-native architectures.
- Strong understanding of data systems (streaming, indexing, caching, ETL).
- Demonstrated experience designing for scalability, fault tolerance, and performance.
Preferred Qualifications
- Experience building RAG systems or embedding-based search.
- Familiarity with vector databases and text/signal retrieval systems.
- Experience developing or deploying agentic AI systems (LLM-based automation, AI observability).
- Strong distributed-systems background (consensus, message buses, job orchestration, eventual consistency).
- Experience designing data schemas and APIs for knowledge representation and operational reasoning.
- Familiarity with ChatOps frameworks or workflow orchestration (Temporal, Argo, Airflow).
- Background in MLOps, AI infrastructure, or platform reliability engineering.
- Experience with observability frameworks (Prometheus, Grafana, OpenTelemetry) and using telemetry for automated reasoning or remediation.
Why CoreWeave
At CoreWeave, we work hard, have fun, and move fast. You’ll join a team that values curiosity, ownership, and creative problem-solving. As part of Production Engineering, you’ll operate at the intersection of AI and reliability — building systems that make operating the most po
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s