Advanced Cooling Facilities Manager
LambdaAbout the role
Lambda, The Superintelligence Cloud, builds Gigawatt-scale AI Factories for Training and Inference. Lambda’s mission is to make compute as ubiquitous as electricity and give every person access to artificial intelligence. One person, one GPU.
If you'd like to build the world's best deep learning cloud, join us.
Travel: 50% , Travel required to various data center sites.
What You’ll Do:
We are seeking an accomplished Advanced Cooling Facilities Manager specializing in Direct Liquid Cooling (DLC) systems to lead the global strategy, implementation, and operational excellence of Lambda’s next-generation liquid cooling infrastructure. This role will define methodologies and standards for the deployment, optimization, and scaling of cooling systems that enable Lambda’s GPU Cloud to deliver industry-leading performance for AI and machine learning workloads.
With deep domain expertise in liquid cooling technologies and critical facilities management, you will drive the design and operation of complex cooling ecosystems, including Coolant Distribution Units (CDUs), hybrid loop architectures, and advanced heat-rejection systems, across colocation and owned data center environments. You will work cross-functionally with internal and external experts to establish best practices, evaluate emerging technologies, and ensure that Lambda’s cooling infrastructure scales reliably and efficiently to support extreme rack densities.
Key Responsibilities:
Liquid Cooling Systems Strategy & Management
CDU Operations & Optimization: Define and oversee operational standards and lifecycle management for all CDU systems (L2L and L2A), including performance optimization, reliability engineering, and capacity expansion strategies. Utilize advanced analytics to identify trends and implement predictive maintenance practices.
Technical Loop Governance: Lead the design and management of multi-stage cooling loops — from facility to rack level — ensuring precise control of temperature, pressure, and flow rate across variable load conditions. Establish system performance benchmarks and quality assurance protocols for coolant integrity and flow balancing.
System Integration Leadership: Coordinate and validate integration of CDUs with facility water systems (FWS), heat exchangers, and mechanical infrastructure. Develop standardized control sequences and commissioning procedures across multiple OEM platforms.
Performance Engineering & Monitoring: Architect the monitoring framework for coolant system telemetry — pressure, temperature, flow, differential, and conductivity — and leverage analytics for continuous improvement in thermal performance, redundancy, and energy efficiency.
Predictive & Preventive Maintenance: Design and institutionalize maintenance methodologies, including condition-based maintenance schedules, failure-mode analysis, and reliability improvement plans for pumps, heat exchangers, and filtration systems.
Infrastructure Planning & Scaling
Capacity Planning & Design Leadership: Evaluate and forecast thermal capacity requirements for high-density GPU clusters, driving design and procurement of CDUs and loop systems to support rack densities exceeding 1 MW. Develop multi-year cooling capacity roadmaps aligned with corporate growth strategies.
Engineering Collaboration: Partner with data center design and mechanical engineering teams to co-develop cooling topologies, redundancy strategies, and modular infrastructure designs optimized for scalability and efficiency.
Vendor & Technology Strategy: Act as the primary technical authority for liquid cooling vendor engagement — influencing product roadmaps, negotiating technical specifications, and qualifying emerging solutions such as direct-to-chip and immersion cooling.
Innovation & Continuous Improvement: Evaluate and pilot next-generation cooling technologies and automation platforms to reduce PUE, enhance reliability, and support sustainability objectives.
Cost & Efficiency Optimization: Establish performance metrics for cooling energy efficiency, uptime, and total cost of ownership. Drive initiatives to reduce CapEx/OpEx through standardization, component reuse, and intelligent control strategies.
Operations & Reliability
Mission-Critical Operations: Oversee global operation of liquid cooling infrastructure with near-zero downtime objectives. Define escalation pro
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s
Similar roles
Advanced Respiratory Therapy Student
Baptist Health
Advanced Practice Provider - Outpatient Orthopedics - General
VCU Health
Advanced Practice Provider - Joint Reconstruction/Sports Medicine - Colonial Heights
VCU Health