Senior HPC Engineer
Oak Ridge National LaboratoryAbout the role
Requisition Id 15468
Overview:
We are hiring a Senior Linux HPC Systems Engineer to design, operate and maintain clusters, servers, and workstations supporting services where science happens at ORNL! This position resides in the Emerging Technologies & Computing team in the Research Computing group in the Information Technology Services Directorate at Oak Ridge National Laboratory (ORNL).
The Emerging Technology Computational Group facilitates ORNL goals through HPC systems engineering, integration, and support for the research community at ORNL. By providing design, deployment, optimization, monitoring, and tooling support across multiple clustered infrastructures, we facilitate Lab-wide R&D projects. Our HPC clusters range in scope from just a handful of nodes to over fifty-thousand cores.
We partner with ORNL research organizations to enable research excellence and delivery. We work with other clustered computing and HPC groups to help research programs identify the best solutions for their needs. When we build our customer's environments, our team collaborates to design, implement, and maintain the systems from inception to retirement.
Major Duties/Responsibilities:
- Advocate and promote HPC and clustered computing services to researchers who process large data sets and/or develop code as a part of their project.
- Ensure the availability, performance, scalability, and security of production systems.
- Leverage automation and monitoring solutions that minimize our day-to-day maintenance and scout opportunities to optimize system management practices or system performance.
- Provide strategic leadership in the design, deployment, and long-term evolution of HPC and clustered computing services, ensuring they align with ORNL’s mission and scientific priorities.
- Set technical direction for HPC infrastructure initiatives, driving innovation in compute, storage, networking, and system management practices to support next-generation scientific workloads.
- Oversee the availability, performance, scalability, and security of mission-critical HPC production systems, balancing operational stability with research agility.
- Lead cross-functional collaborations with program POCs, research teams, and other HPC groups to architect solutions, optimize performance, and accelerate scientific discovery.
- Champion automation and advanced monitoring frameworks that reduce operational overhead, improve reliability, and provide predictive insights into system performance.
- Mentor and guide junior engineers and administrators, fostering professional growth and ensuring knowledge transfer within the HPC operations team.
- Represent ORNL in external collaborations, engaging with vendors, DOE, and peer institutions to influence technology roadmaps and adopt emerging HPC innovations.
- Develop and enforce best practices and standards for system deployment, monitoring, security, and lifecycle management across multiple HPC environments.
- Contribute to long-term planning, including capacity forecasting, technology evaluation, and adoption of advanced computing paradigms
- Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace – in how we treat one another, work together, and measure success.
Basic Qualifications:
- A BS degree in computer science, computer engineering, information technology, information systems, science, engineering, business, or a related discipline and a minimum of eight (8) to twelve (12) years of aligned professional experience is required for consideration. An overall combination of equivalent education and experience may be considered.
- Masters and PhD degree holders in the same fields of study are also encouraged to apply:
-
-
- Masters’ holders should have a minimum of seven (7) to ten (10) years of relevant and aligned experience.
- PhD holders should have a minimum of four (4) to six (6) years of relevant and aligned experience.
-
- Five (5) or more years with managing UNIX/Linux Systems.
- Three (3) or years of proven experience with configuration management and automation
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s