Jobs and Careers
NE
Senior HPC Cluster Engineer
NebiusCzechiaRemotefull_timeVerifiedPosted 29 May 2026
About the role
<div><p><strong>About Nebius:</strong></p>
<p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p>
<p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p>
<p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.</p></div><h3><strong><span>The role</span></strong></h3>
<p>We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.</p>
<p> </p>
<p><strong>In this position, you will be responsible for:</strong></p>
<ul>
<li><strong>Tuning the performance</strong> of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments.
</li>
<li><strong>Analyzing and troubleshooting</strong> the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions.
</li>
<li><strong>Integrating new hardware</strong> into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM.
</li>
<li><strong>Enhancing automation systems </strong>for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments.
</li>
<li><strong>Configuring and managing GPU devices</strong> and InfiniBand fabrics, ensuring efficient and reliable operation.<strong><br/></strong></li>
</ul>
<p> </p>
<p><strong>We expect you to have:</strong></p>
<ul>
<li>5+ years of professional experience in <strong>system-level software development</strong> (focused on performance optimization, low-level programming).
</li>
<li>3+ years of hands-on experience with <strong>Linux systems</strong> (administration, troubleshooting, and performance tuning).
</li>
<li><strong>In-depth understanding</strong> of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.
</li>
<li>Strong proficiency in one or more <strong>performance-oriented programming languages</strong> (C/C++, Go, Python).</li>
</ul>
<p> </p>
<p><strong>It would be a plus if you have:</strong></p>
<ul>
<li>Experience with <strong>GPU end-to-end testing</strong> in a <strong>cluster environment</strong> using InfiniBand networking.
</li>
<li>Proven track record of analyzing and optimizing the performance of <strong>HPC workloads</strong> (e.g., simulations, data analysis, AI/ML workloads).
</li>
<li>Familiarity with <strong>RDMA, RoCE, and InfiniBand</strong> protocols for high-performance communication.
</li>
<li>Background in <strong>Software-Defined Networking</strong> (SDN) and experience with <strong>HPC cluster networking</strong>.
</li>
<li>Understanding of <strong>QEMU/KVM virtualization</strong> and managing virtualized environments.
</li>
<li>Experience with <strong>deep learning frameworks</strong> such as <strong>PyTorch</strong> and <strong>TensorFlow</strong>, and their integration with HPC systems.
</li>
<li>Familiarity with <strong>collective communication libraries</strong> like <strong>MPI</strong> and <strong>NCCL</strong> for distributed computing. </li>
</ul>
<p><span><em>We conduct coding interviews as part of the process.</em></span></p><div><p><strong>Benefits & Perks:</strong></p>
<ul>
<li>Competitive compensation</li>
<li>Career growth and learning opportunities</li>
<li>Flexibility and work-life balance</li>
<li>Collaborative and innovative culture</li>
<li>Opportunity to work on impactful AI projects</li>
<li>International environment and talented teams</li>
</ul>
<p><strong>What's it like to work at Nebius:</strong></p>
<p>Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI </p>
<p><strong>Equal Opportunity Statement:</strong></p>
<p>Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment o
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s