Jobs and Careers
CO

Staff Software Engineer, Compute Architecture

CoreWeave
USAfull_timePosted 1 Jul 2026

About the role

<div class="content-intro"><div> <div> <div class="gmail_quote"> <div> <div><span id="m_1770241969069985273m_-2746164444908759431gmail-docs-internal-guid-131e4fb0-7fff-b4e9-ff50-e8cf32449b1b">CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at&nbsp;<a href="http://www.coreweave.com/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=http://www.coreweave.com&amp;source=gmail&amp;ust=1762613132717000&amp;usg=AOvVaw3D-UOhNaqEvF5BEWxjYyAU">www.coreweave.com</a>.</span></div> </div> </div> </div> </div></div><h4><strong>About the Role</strong></h4> <p>As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers. The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.</p> <h4><strong>What You’ll Do</strong></h4> <ul> <li>Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.</li> <li>Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.</li> <li>Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.</li> <li>Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly.</li> <li>Translate incidents and hardware failure modes into software improvements that make the platform more resilient.</li> <li>Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale.</li> <li>Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship.</li> <li>Make pragmatic architectur

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

CoreWeave

View company profile →