Jobs and Careers
TE

HPC System Engineer

TerraPower
United Statesfull_timeVerifiedPosted 7 Apr 2025
💰 $170,408/yr($113,605/yr$170,408/yr)

About the role

Location: Bellevue,Washington,United States

TITLE:  HPC System Engineer

LOCATION:     Bellevue, WA

TerraPower is a nuclear technology company based in Bellevue, Washington. At its core, the company is working to raise living standards globally through a more affordable, secure and environmentally friendly form of nuclear energy along with innovations in medical isotopes to improve human health. In 2006, TerraPower originated with Bill Gates and a group of like-minded visionaries who evaluated the fundamental challenges to raising living standards around the world. They recognized energy access was crucial to the health and economic well-being of communities and decided that the private sector needed to take action and create energy sources that would advance global energy deployment. TerraPower’s mission is to be a world leader in new nuclear technologies, while developing innovators and future leaders in the nuclear field. As a result, the company’s activities in the fields of nuclear energy and related sciences are yielding significant innovations in the safety and economics of nuclear power, hybrid energy and medical applications – all for significant human health benefits.

TerraPower is seeking to hire highly motivated and forward-thinking professionals who are interested in focusing on advanced nuclear reactor research and development and influencing change within the nuclear power landscape and bringing forward the critical production of medical isotopes.  TerraPower is an Equal Opportunity Employer. We do not discriminate in hiring on the basis of sex, gender identity, sexual orientation, race, color, religious creed, national origin, physical or mental disability, protected Veteran status, or any other characteristic protected by federal, state, or local law. In addition, as a federal contractor, TerraPower has instituted an Affirmative Action Plan (AAP) in an effort to proactively recruit, hire, and promote women, minorities, disabled persons and veterans.

HPC System Engineer

The HPC (High-Performance Computing) System Engineer is responsible for partnering with internal customers who are leveraging complex advanced physics and engineering applications with the goal to provide a best-in-class end-user experience through resilient, capable platforms and solutions. Role will supply technical input to the infrastructure requirements and provide technical support, vendor management and training to users of these resources for advanced nuclear. This position requires an experienced candidate who will promote collaboration and cooperation while working with multiple engineering disciplines.

Responsibilities:

    Maintain Linux HPC Supercomputer systems availability to the customer, including in Azure Gov and on-prem infrastructure.

    Administer and maintain Linux based system software and firmware revisions, including patches, updates, and OS upgrades.

    Solve Linux system hardware, software, and third-party software issues, and provide detailed and thoughtful analysis of problem and resolution.

    Automate configuration management of infrastructure and applications, software updates, and maintenance and monitoring of system availability using modern DevOps tools (Ansible, GitHub, etc.)

    Installation, configuration, tuning, troubleshooting, and administration of commercial off-the-shelf (COTS), Open Source, and in-house developed applications leveraging HPC resources.

    Packaging, deployment, and management of software leveraging environment modules.

    Coordinate HPC infrastructure solutions and plan for growth.

    Actively connect with management regarding any problems with the equipment and propose resolution.

    Partner with IT Principal Engineering to define and execute roadmaps. Assist with gathering data for new feature, system, and/or advanced computing requirements from key stakeholders.  Provide timely estimates for implementation delivery. Anticipate risks when planning and defining mitigation options.

    Respond to user queries regarding computing resources.

Key Qualifications and Skills

    BA, BS, or MS in CS, EE, CE or equivalent experience.

    5+ years of previous experience deploying and administrating production HPC clusters.

    Experience with managing an HPC resource scheduler (Slurm preferred).

    Proven track record to script in Bash or Python.

    Experience with MPI software and high-speed interconnects in HPC supercomputers.

    Experience with containe

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

TerraPower

View company profile →