Senior High Performance Computing Engineer - Classified Environment
Oak Ridge National LaboratoryAbout the role
Requisition Id 14691
Overview:
The Field Intelligence Operations Division is seeking a Senior High Performance Computing (HPC) Engineer for Classified Computing to lead the design, implementation, and management of HPC systems within a classified environment. We are looking for candidates with extensive experience in HPC architecture, cluster management, and parallel computing, with a proven ability to work within highly secure and regulated environments. This role involves close collaboration with security teams, scientists, and IT leadership to ensure that the HPC infrastructure meets the stringent performance, security, and compliance requirements necessary for classified work.
As part of our team, you will join a dynamic and elite group of professionals specializing in the design, implementation, and management of HPC systems to support cutting-edge computational needs. Our team is highly collaborative, striving to ensure a gold standard in HPC architecture and operations are understood, implemented, and optimized for performance. You will play a critical role in delivering exceptional service to users and stakeholders by supporting system deployment, configuration, training, and education, ensuring that HPC resources are accessible, secure, and operating at peak efficiency.
Major Duties/Responsibilities:
- HPC System Design and Architecture:
- Lead the design and deployment of HPC systems, ensuring they meet the computational needs and security requirements of a classified environment.
- Create and maintain detailed documentation of HPC architectures, configurations, and operational procedures.
- Cluster Management and Optimization:
- Oversee the installation, configuration, and management of HPC clusters, ensuring optimal performance, scalability, and reliability.
- Implement and manage job scheduling, resource allocation, and load balancing to maximize the efficiency of HPC resources.
- Security and Compliance:
- Ensure all HPC systems comply with security policies and regulatory requirements, implementing necessary controls and conducting regular audits.
- Collaborate with the security team to address vulnerabilities and ensure the protection of sensitive data within the HPC environment.
- Performance Tuning and Troubleshooting:
- Monitor and optimize the performance of HPC systems, identifying and resolving bottlenecks and inefficiencies.
- Identify and resolve complex issues, ensuring minimal downtime and disruption to critical operations.
- Collaboration and Leadership:
- Lead HPC-related projects, from initial planning and design through to implementation and operational support.
- Collaborate with scientists, researchers, and others to ensure that the HPC environment meets their computational needs.
- Mentor and support junior HPC engineers, sharing expertise and best practices.
- Continuous Improvement and Innovation:
- Research and remain informed of the latest advancements in HPC technologies, identifying opportunities for innovation and enhancement of the HPC infrastructure.
- Propose and implement improvements to existing systems and processes to support the evolving needs of the organization.
Basic Qualifications:
- BS in computer science, engineering, or a related field and eight (8) years of relevant experience. An equivalent combination of education and experience may be considered.
- Seven (7) years of experience in HPC engineering, with a focus on cluster management, parallel computing, and performance optimization.
- Demonstrated experience working in classified environments, i
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s