Jobs and Careers
CO

Infrastructure Reliability Engineer, Bare Metal

CoreWeave
United StatesRemotefull_timeVerifiedPosted 4 Nov 2025
💰 $163,000/yr($122,000/yr$163,000/yr)

About the role

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

We seek a highly skilled and driven Infrastructure Reliability Engineer, Bare Metal to join our team and report to our Senior Director, Customer Experience. In this role, you will be instrumental in ensuring the stability, performance, and ongoing improvement of our intricate bare metal infrastructure. This role demands a deep technical understanding of underlying hardware and related systems, coupled with a proactive approach to problem-solving and operational efficiency. You will collaborate extensively with diverse engineering teams and external vendors, driving automation initiatives and refining operational strategies within a rapidly expanding, technologically advanced environment. Your contributions will directly impact the core of our compute capabilities.

Please note: This role will be based in Pacific Time Zone (PST) hours.

What You'll Do

  • Provide expert-level technical support and in-depth troubleshooting for a wide spectrum of hardware and associated software issues, encompassing server malfunctions, network outages, and performance degradations.
  • Manage the lifecycle of our bare metal infrastructure, including overseeing deployment methodologies, executing maintenance procedures, coordinating upgrades, and managing hardware retirement processes.
  • Architect and implement automation solutions through scripting and tooling to streamline repetitive operational tasks, enhance overall efficiency, and minimize manual intervention across the infrastructure.
  • Lead the development and refinement of critical operational processes, comprehensive technical documentation (SOPs, TSGs, runbooks), and the establishment of engineering best practices to bolster team effectiveness and infrastructure resilience.
  • Engage in close collaboration with Software, Network, and Data Center Operations Engineering teams to facilitate effective issue resolution, contribute to strategic project planning, and ensure the cohesive operation of the entire infrastructure ecosystem.
  • Serve as a key technical point of contact for hardware and software vendors, managing technical support engagements, overseeing the RMA process, and driving the resolution of complex hardware-centric challenges.
  • Design, deploy, and maintain sophisticated monitoring and alerting frameworks to proactively identify and mitigate potential infrastructure anomalies and performance deviations.
  • Participate actively in incident response protocols, conduct thorough root cause analysis (RCAs) for infrastructure events, and contribute to problem management strategies aimed at preventing future occurrences.
  • Contribute technical expertise to and potentially lead infrastructure-focused projects, including new hardware deployments, critical system upgrades, and the integration of new operational tooling.
  • Mentor and guide junior engineering team members, fostering technical growth and contributing to the development of internal knowledge resources and training programs.
  • Maintain the integrity of hardware asset tracking and related data within our infrastructure inventory systems (e.g., Snipe-IT).
  • Adhere to and promote stringent security protocols and best practices related to infrastructure access and maintenance activities.

Who You Are

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent experience
  • 5+ years of experience in hands-on management and support of complex bare metal infrastructure environments and data center operations
  • Comprehensive understanding of modern server hardware architectures, including specialized compute accelerators (GPUs) and high-speed interconnect technologies from leading high-performance computing vendors such as NVIDIA, Dell, or HPE.
  • Demonstrated expertise in Linux system administration, encompassing deep familiarity with command-line operations and system configuration.
  • Proficiency in at least one high-level scripting language (e.g., Python) and practical experience with infrastructure and/or network automation tools, methodologies, and frameworks (

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

CoreWeave

View company profile →