Jobs and Careers
AH

HPC Infrastructure Engineer

AHEAD
United States, United Statesfull_timeVerifiedPosted 21 Jan 2025

About the role

AHEAD builds platforms for digital business. By weaving together advances in cloud infrastructure, automation and analytics, and software delivery, we help enterprises deliver on the promise of digital transformation.
At AHEAD, we prioritize creating a culture of belonging, where all perspectives and voices are represented, valued, respected, and heard. We create spaces to empower everyone to speak up, make change, and drive the culture at AHEAD. 
We are an equal opportunity employer, and do not discriminate based on an individual's race, national origin, color, gender, gender identity, gender expression, sexual orientation, religion, age, disability, marital status, or any other protected characteristic under applicable law, whether actual or perceived. 
We embrace all candidates that will contribute to the diversification and enrichment of ideas and perspectives at AHEAD. 
The High-Performance Computing Infrastructure Engineer is primarily responsible for the overall health and maintenance of HPC infrastructure in our managed services customer's environments. Our HPC Infrastructure Engineers are a valued member of the Managed Services Infrastructure Practice responsible for Tier 3 incident management, service request management and change management infrastructure support for all Managed Services customers. 

Roles & Responsibilities

  • Providing enterprise-level operational support to Managed Services customers for incident, problem, and change management activities 
  • Design, deploy, and manage Kubernetes clusters optimized for HPC workloads, with a focus on integrating and managing NVIDIA DGX systems. 
  • Optimize cluster performance, resource utilization, and cost-effectiveness, specifically addressing the unique requirements of DGX systems. 
  • Implement monitoring, logging, and alerting solutions for HPC Linux clusters, Kubernetes, and DGX infrastructure 
  • Ensure the security of the Kubernetes infrastructure and HPC workloads, including the protection of sensitive data processed by DGX systems. 
  • Troubleshoot and resolve issues related to Kubernetes, DGX systems, HPC applications, and infrastructure 
  • Stay up to date on the latest technologies and trends in Kubernetes, HPC, and NVIDIA DGX systems, including new hardware and software releases 
  • Work across technical teams to troubleshoot complex infrastructure issues 
  • Create and maintain detailed documentation 
  • Serve as a subject matter expert and escalation point for HPC technologies 
  • Work with vendors to resolve infrastructure issues 
  • Communicate with customers and internal team with transparency 
  • Participate in on-call rotation 
  • Completion of training and certification as assigned to further skills and knowledge 

Qualifications

  • Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education
  • 5+ years of expert level experience managing infrastructure in high-performance computing environments including configuration, troubleshooting, and best practice
  • Strong understanding of Kubernetes architecture, components, and networking
  • Hands-on experience with deploying, managing, and optimizing NVIDIA DGX systems preferred
  • Linux engineer with experience in RedHat, Ubuntu, and Rocky distributions
  • Experience with deploying and managing Kubernetes clusters in production environments, including those with GPU acceleration
  • Experience with HPC workloads, schedulers (e.g., SLURM, PBS, Torque), and applications, particularly in the context of AI/ML and deep learning
  • Experience with containerization technologies (e.g., Docker, Singularity)
  • Experience with Infrastructure-as-Code (IaC) tools (e.g., Terraform, Ansible)
  • Experience with monitoring and logging tools (e.g., Prometheus, Grafana), experience integrating with Elastic Observability
  • Strong scripting skills (e.g., Bash, Python)
  • Excellent problem-solving and troubleshooting skills
  • Experience configuring, maintaining and troubleshooting Kubernetes
  • Experience with storage technology (e.g., Ceph, Vast Data Platform) and distributed file systems (e.g., Lustre, GPFS, NFS, GlusterFS)
  • Experience configuring, maintaining and troubleshooting Nvidia/Mellanox (Cumulus OS) switches a plus
  • Experience with both ethernet and InfiniBand networking a plus
  • 1+ years working with an enterprise ITSM system: Service Now is a bonus
  • Managed Services or consulting experience is required
  • Strong background with customer service
  • High level problem-solving and communication skills
  • Strong oral and written communications skills
  • Related certific

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

AHEAD

View company profile →