Jobs and Careers
CA

IT InfiniBand/GPU -Sr Staff Systems Engineer

Cadence
San Jose, United Statesfull_timeVerifiedPosted 10 Jun 2024
💰 $247,000/yr($133,000/yr$247,000/yr)

About the role

At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.

Cadence is looking for a Sr Staff Systems Engineer who accelerates strategic customer deployments and ensures on-time bring-up and deployment of HPC infrastructure and troubleshooting and supports technical roles supporting HPC, InfiniBand, and GPU at our San Jose location!

The successful candidate will be a hands-on technical candidate within the infrastructure team and be exposed to customer interfaces dealing with the Windows and Linux OS.

The System Engineer will need experience in Linux environments and proficiency in tasks such as shell scripting.

Role: IT -Sr Staff Systems Engineer

Location on-site (not remote): San Jose, CA

Must Haves

  • 15+ years of experience in system administration and engineering.

  • Minimum five years overall experience in technical roles supporting GPU Infrastructure setup using InfiniBand

  • Experience with interconnections between InfiniBand & GPU’s

  • Experience with GPU Enabled MPI’s

  • Experience with GPU Nvidia CUDA or AMD’s ROCm

  • Experience with; H100, AMD MI210, GPU servers in Cluster   

  • Customer deployments and ensure on-time bring-up of GPU Servers.  InfiniBand fabric bring-up, configuration, and subnet management on the IB switch

  • Participate in engagements with various SW and FW (BMC/SBIOS/OS/drivers etc.) teams to develop best-in-class practices and tools; you will be analyzing, debugging, and resolving critical firmware and software issues for the workload performance at scale

  • Provide engineering solutions to enable large-scale performance strategies for performance for Datacenter GPU Computing products and software stacks, ensure technical relationships with internal and external engineering teams, and assist systems engineers in building creative solutions

  • Strong knowledge of Linux operating systems and networking and security concepts.

  • Document and drive acceptance and qualification test plans, procedures, and reports

Requirements

  • Accelerate strategic customer deployments and ensure on-time bring-up and deployment of HPC infrastructure

  • Participate in engagements with various SW and FW (BMC/SBIOS/OS/drivers etc.) teams to develop best-in-class practices and tools; you will be analyzing, debugging, and resolving critical firmware and software issues for the workload performance at scale

  • Provide engineering solutions to enable large-scale performance strategies for performance for Datacenter GPU Computing products and software stacks, ensure technical relationships with internal and external engineering teams, and assist systems engineers in building creative solutions

  • Development and implementation of server and rack-level telemetry aspects, collaborate and establish continuous improvements in our design flows

  • Recent experience in critical data center technologies such as server architectures, software containers, job schedulers, and parallel computing. Deployment and operation of large-scale systems; resilient system design; and clustering of computing resources

  • cluster management for HPC and actively connect with management regarding any problems with the equipment and propose a resolution

  • Establish and maintain IT infrastructure and procedures for customer-facing and internal systems

  • Actively establish the technical relationship with our customer’s engineers, management, and architects at focus accounts

  • Create and develop test plans for new features on each product. Recommend improvements to enable automated scripting for testing and archiving of results. Develop HPC computing strategies for cloud-based computing, GPU-accelerated computing, etc.

  • Provide remote cluster support to large environments, including scalability/flexibility and troubleshooting end-user issues involving job submission, runtime, and resource access.

  • InfiniBand fabric configuration and administration on Red hat/Centos/Linux experience in configuring PKeys and troubleshooting the end-to-end InfiniBand environment

  • InfiniBand fabric bring-up, configuration, subnet management, and monitoring on the IB switch and client side for multi-tenancy setup, understanding of IPoIB communication modes

  • Performance comparison of the InfiniBand network with cluster interconnects and debugging the InfiniBand performance-related issues

  • Automate configuration management, software updates, and system availability maintenance and monitoring using modern DevOps tools (Ansible, Gitlab, etc.)

  • Be a technical specialist on GPU computing and network

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Cadence

View company profile →