Technical Expert/Functional Expert (HPC)
Cyber B.A.T. Inc.About the role
Join the Mission at Cyber B.A.T.
At Cyber B.A.T., we believe in one simple idea: Take care of those that take care of you.
We’re a Service-Disabled Veteran-Owned Small Business delivering engineering, cybersecurity, software, systems, and mission-focused consulting solutions to government customers and industry partners. We combine the impact and innovation of a growing company with the culture and flexibility of a close-knit team.
When you join Cyber B.A.T., you’re not just taking a job, you’re helping solve meaningful challenges alongside people who value collaboration, technical excellence, and having fun while doing it.
Why People Love Working Here
- Employee ownership opportunities — every employee has equity in our success
- Annual company retreats (yes… actual destinations)
- 5 weeks PTO + flexible work environment
- Employer-paid medical and dental for employees and dependents
- Up to 10% 401(k)/profit sharing contribution (no match required)
- Monthly happy hours and team events
- Referral bonuses
- Opportunity to make visible impact without getting lost in a giant organization
Position Overview
Cyber B.A.T. is seeking a Technical Expert / Functional Expert – High Performance Computing (HPC) to provide Tier 4 technical support for enterprise High Performance Computing (HPC) environments supporting mission-critical operations. This role serves as a senior subject matter expert responsible for the configuration, optimization, troubleshooting, and sustainment of large-scale HPC infrastructure, distributed computing environments, and high-performance storage systems.
What You'll Do
- Provide Tier 4 operational support for enterprise High Performance Computing (HPC) environments supporting mission-critical workloads
- Configure, optimize, test, and troubleshoot high-performance file systems, including XFS, GPFS, and Lustre
- Administer, tune, and maintain distributed computing and workload management platforms, including RES, LSF, and SLURM
- Perform advanced troubleshooting and performance optimization of HPC clusters, compute farms, and associated applications
- Configure, maintain, and support HPC management and monitoring tools, including Nagios, xCAT, failover solutions, and compiler toolchains
- Support Massively Parallel Processing (MPP) systems and distributed computing architectures
- Administer, configure, tune, and troubleshoot Red Hat Enterprise Linux (RHEL) and SUSE Linux Enterprise Server (SLES) operating systems
- Monitor system health, availability, capacity, and performance to ensure continuous mission operations
- Perform root cause analysis for complex infrastructure, storage, operating system, and distributed computing issues
- Develop and implement performance tuning recommendations to improve system scalability, reliability, and efficiency
- Collaborate with systems engineers, software engineers, storage administrators, and infrastructure teams to support enterprise HPC environments
- Develop and maintain technical documentation, operational procedures, and engineering best practices
- Provide technical guidance and subject matter expertise to engineering teams supporting HPC operations and modernization initiatives
- Evaluate and recommend emerging HPC technologies, tools, and operational improvements to enhance mission capabilities
What We're Looking For
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Systems Engineering, or a related technical discipline
- Minimum of 10 years of experience supporting large, complex enterprise IT or High Performance Computing (HPC) environments; or 15 years of directly related professional experience in lieu of a bachelor's degree
- Demonstrated expertise configuring, tuning, testing, and troubleshooting enterprise HPC environments
- Experience administering and optimizing high-performance file systems, including: XFS, GPFS (IBM Spectrum Scale), Lustre
- Experience administering distributed computing and workload scheduling platforms, including: RES, LSF, SLURM
- Experience supporting HPC cluster management tools and infrastructure
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s