Jobs and Careers
GU

Senior HPC Infrastructure Engineer

Guardant Health
Palo Alto, United Statesfull_timeVerifiedPosted 24 Sept 2024
💰 $187,300/yr($138,700/yr$187,300/yr)

About the role

Company Description

Guardant Health is a leading precision oncology company focused on helping conquer cancer globally through use of its proprietary tests, vast data sets and advanced analytics. The Guardant Health oncology platform leverages capabilities to drive commercial adoption, improve patient clinical outcomes and lower healthcare costs across all stages of the cancer care continuum. Guardant Health has commercially launched Guardant360®, Guardant360 CDx, Guardant360 TissueNext™, Guardant360 Response™, and GuardantOMNI® tests for advanced stage cancer patients, and Guardant Reveal™ for early-stage cancer patients. The Guardant Health screening portfolio, including the Shield™ test, aims to address the needs of individuals eligible for cancer screening.

Job Description

Guardant’s HPC team builds and operates the computational technology backbone of the company. 

This includes scalable data storage that holds PBs of genomics data, high performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration. To facilitate Guardant Health’s fast growth in the next few years, the HPC team is looking for a strong technical engineer who can help maintain and help grow the HPC infrastructure during its aggressive expansion, while working with corporate IT, SQA and DevOps/SRE teams. 

This role can be remotely worked part-time, but requires a very hands on, on-premise presence when on rotation, minimally.

 In this role, you will primarily:

  • Assist in managing the HPC interconnect
  • Assist in integrating the HPC systems with the bandwidth on-demand system
  • Work with the networking infrastructure team to manage and optimize the connectivity to and from the HPC systems and locales
  • Help manage multiple HPC clusters and cluster file systems. 
  • Help research, develop and implement the next generation HPC solution
  • Troubleshoot the production system stack down to source code level e.g. shell scripts, python and others.
  • Maintain, monitor, and support the infrastructure environment and/or facilities.
  • Use and maintain enhanced production monitoring and additional capability.
  • Support improvements for increased system reliability and performance.
  • Support multiple systems or applications of medium to high complex (complexity defined by size, technology used, and system feeds and interfaces) with multiple concurrent users, ensuring control, integrity, and accessibility.
  • Support systems at remote locations, including internationally
  • Work with offsite consultants to maintain the infrastructure
  • Work with vendors to troubleshoot, upgrade and repair systems as needed
  • Participate in a 24/7 on-call rotation

Qualifications

You enjoy an agile, very fast paced and highly technical environment. You are a self-driven accomplished technologist who strives to be ever improving your skills, value to the company and improve the computational infrastructure.  You are dedicated to engineering excellence yet pragmatic and flexible.  You have the ability to maintain the day-to-day support SLA while running various key projects that move the business forward. 

  • 2+ years of Linux/Unix administration, knowledge of Unix network protocols, TCP/IP network fundamentals, core infrastructure technologies and virtualization 
  • 2+ years of large-scale data storage and compute clusters (HPC) infrastructure  
  • 2+ years working in and with on-premise and cloud-based (AWS, Google, IBM and Azure) data-centers 
  • 2+ years of building software release and ops processes and automation toolset 
  • 2+ years providing documentation of system administration

Following Skills Sets are Preferred:

  • Experience administering IBM’s General Parallel File System
  • Experience administering Grid Engine scheduler
  • Experience administering SLURM scheduler
  • Experience with using Bright Cluster Manager
  • Experience with cloud bursting technologies
  • Experience with wide area file systems
  • Experience with docker and container technologies
  • Experience with Kubernetes, preferably with Certified Kubernetes Administrator (CKA)
  • Operating infrastructure compliant with HIPAA and SOX standards

Education

B.S. in Computer Science or related field

Additional Information

Hybrid Work Model: At Guardant Health, we have defined days for in-person/onsite collaboration and work-from-home days for individual-focused time. All U.S. employees who live within 50 miles of a Guardant facility will be required to be onsite on Mondays, Tuesdays, and Thu

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Guardant Health

View company profile →