HPC Engineer
The Aerospace CorporationAbout the role
The Aerospace Corporation is the trusted partner to the nation’s space programs, solving the hardest problems and providing unmatched technical expertise. As the operator of a federally funded research and development center (FFRDC), we are broadly engaged across all aspects of space— delivering innovative solutions that span satellite, launch, ground, and cyber systems for defense, civil and commercial customers. When you join our team, you’ll be part of a special collection of problem solvers, thought leaders, and innovators. Join us and take your place in space.
Job Summary
The Aerospace Corporation is seeking a talented and motivated High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will be responsible for developing, implementing, and optimizing HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges and contributing directly to the success of critical national space assets. We value a collaborative, proactive mindset and a shared commitment to engineering excellence.
Work Model
This is a full-time position located in either Chantilly, VA, or El Segundo, CA, with an expectation of 100% onsite work.
What You’ll Be Doing
- Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings.
- Manage a 10,000-core classified cluster and a 3,000-core unclassified cluster to ensure peak performance.
- Deliver high-quality HPC infrastructure design, automated provisioning, and system configuration.
- Develop and deploy automation solutions using tools such as Ansible or Puppet.
- Optimize AI workloads and GPU computing performance.
- Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency.
- Work closely with scientists and engineers to support new and ongoing projects, and mission technical analysis supporting national space assets.
- Develop cost-efficient on-premise and cloud HPC service offerings that align with mission and business objectives.
- Implement and enforce security best practices that comply with government regulations across both classified and unclassified environments.
- Harden Linux systems to meet stringent security requirements
What You Need to be Successful
Minimum Requirements for the Site Reliability Engineer Staff III:
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- Minimum of 5 years’ experience in Linux system administration within an enterprise HPC environment.
- In-depth knowledge of Linux, networking, and HPC systems.
- Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.
- Proficiency in scripting and competence with automation tools such as Ansible or Puppet.
- Experience hardening Linux systems to meet security requirements
- Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco.
- Strong communication skills, with an ability to work both independently and as part of a geographically distributed team.
- CompTIA Security+ CE certification or equivalent that meets DoD 8570.01-m requirements for IAT Level II personnel
- Active TS/SCI clearance. U.S citizenship is required to obtain security clearance.
In addition to the above, the minimum requirements for the Site Reliability Engineer Staff IV include:
- 7+ years of experience in an enterprise HPC environment.
- Experience performing in-place upgrades of Slurm.
- Experience provisioning and supporting AI & NVIDIA GPU technologies
- Skill in provisioning and supporting AI & NVIDIA GPU technologies, with expertise in GPU integration, resource allocation, and scheduling using Slurm.
- Hands-on background with cloud HPC services
- Experience supporting a wide range of technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules.
- Demonstrated ability to lead cross-functional teams and mentor junior engineers.
How You Can Stand Out
It would be impressive if you have one or more of these:
- Experience with AWS Parallel Computing Service (AWS ParallelCluster).
- Knowledge of NVLINK or NVSWITCH for optimizing GPU workflows.
- Familiarity with Prometheus and Grafana for monitoring and performance visualization.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s