Jobs and Careers
LA

Advanced Cooling Facilities Manager

Lambda
Remote, USA, United StatesRemotefull_timeVerifiedPosted 17 Oct 2025
💰 $204,000/yr($131,000/yr$204,000/yr)

About the role

Lambda, The Superintelligence Cloud, builds Gigawatt-scale AI Factories for Training and Inference. Lambda’s mission is to make compute as ubiquitous as electricity and give every person access to artificial intelligence. One person, one GPU.


If you'd like to build the world's best deep learning cloud, join us. 


Travel: 50% , Travel required to various data center sites.


What You’ll Do:

We are seeking an accomplished Advanced Cooling Facilities Manager specializing in Direct Liquid Cooling (DLC) systems to lead the global strategy, implementation, and operational excellence of Lambda’s next-generation liquid cooling infrastructure. This role will define methodologies and standards for the deployment, optimization, and scaling of cooling systems that enable Lambda’s GPU Cloud to deliver industry-leading performance for AI and machine learning workloads.

With deep domain expertise in liquid cooling technologies and critical facilities management, you will drive the design and operation of complex cooling ecosystems, including Coolant Distribution Units (CDUs), hybrid loop architectures, and advanced heat-rejection systems, across colocation and owned data center environments. You will work cross-functionally with internal and external experts to establish best practices, evaluate emerging technologies, and ensure that Lambda’s cooling infrastructure scales reliably and efficiently to support extreme rack densities.

Key Responsibilities:

Liquid Cooling Systems Strategy & Management

  • CDU Operations & Optimization: Define and oversee operational standards and lifecycle management for all CDU systems (L2L and L2A), including performance optimization, reliability engineering, and capacity expansion strategies. Utilize advanced analytics to identify trends and implement predictive maintenance practices.

  • Technical Loop Governance: Lead the design and management of multi-stage cooling loops — from facility to rack level — ensuring precise control of temperature, pressure, and flow rate across variable load conditions. Establish system performance benchmarks and quality assurance protocols for coolant integrity and flow balancing.

  • System Integration Leadership: Coordinate and validate integration of CDUs with facility water systems (FWS), heat exchangers, and mechanical infrastructure. Develop standardized control sequences and commissioning procedures across multiple OEM platforms.

  • Performance Engineering & Monitoring: Architect the monitoring framework for coolant system telemetry — pressure, temperature, flow, differential, and conductivity — and leverage analytics for continuous improvement in thermal performance, redundancy, and energy efficiency.

  • Predictive & Preventive Maintenance: Design and institutionalize maintenance methodologies, including condition-based maintenance schedules, failure-mode analysis, and reliability improvement plans for pumps, heat exchangers, and filtration systems.

Infrastructure Planning & Scaling

  • Capacity Planning & Design Leadership: Evaluate and forecast thermal capacity requirements for high-density GPU clusters, driving design and procurement of CDUs and loop systems to support rack densities exceeding 1 MW. Develop multi-year cooling capacity roadmaps aligned with corporate growth strategies.

  • Engineering Collaboration: Partner with data center design and mechanical engineering teams to co-develop cooling topologies, redundancy strategies, and modular infrastructure designs optimized for scalability and efficiency.

  • Vendor & Technology Strategy: Act as the primary technical authority for liquid cooling vendor engagement — influencing product roadmaps, negotiating technical specifications, and qualifying emerging solutions such as direct-to-chip and immersion cooling.

  • Innovation & Continuous Improvement: Evaluate and pilot next-generation cooling technologies and automation platforms to reduce PUE, enhance reliability, and support sustainability objectives.

  • Cost & Efficiency Optimization: Establish performance metrics for cooling energy efficiency, uptime, and total cost of ownership. Drive initiatives to reduce CapEx/OpEx through standardization, component reuse, and intelligent control strategies.

Operations & Reliability

  • Mission-Critical Operations: Oversee global operation of liquid cooling infrastructure with near-zero downtime objectives. Define escalation pro

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Lambda

View company profile →