Jobs and Careers
TH

Senior Engineering Manager, AI Infrastructure

The Allen Institute for AI
United Statesfull_timeVerifiedPosted 9 Jul 2026
💰 $220,320/yr($146,880/yr$220,320/yr)

About the role

Persons in these roles are expected to work from our offices in Seattle. On-site requirements vary based on position and team. If you have questions about on-site work arrangements for this role, please ask your recruiter.
Our base salary range is $146,880 - $220,320, and in addition we have generous bonus plans to provide a competitive compensation package. 

Who You Are:

We are seeking a Senior Manager, AI Infrastructure to run the day-to-day operation of the systems that power our research. Reporting to the VP of Engineering, you will own the execution and reliability of our high-performance computing (HPC) environment which includes on-prem GPU clusters and the software orchestration layer that schedules workloads across a hybrid cloud environment. This is a hands-on operational leadership role: your mandate is to keep the platform fast, reliable, and well-utilized, and to deliver against the roadmap set with your PM counterpart.

Our ideal candidate is a:

  • Systems Expert: You have a deep, hands-on understanding of the Linux kernel, container runtimes, and distributed systems. You understand the performance implications of InfiniBand topologies and NCCL optimizations.
  • Execution-Focused Leader: You plan and deliver against near-term operational goals, keep reliability and researcher velocity high, and turn priorities set with leadership into shipped, dependable systems.
  • Pragmatic Operator: You are comfortable making trade-offs between technical elegance and operational necessity. You triage and mitigate immediate risks, and know when to handle something yourself versus escalate.

Who We Are: 

Ai2 is a non-profit research institute at the forefront of open-source AI development. Unlike industry peers, our goal is to share our findings, data, code, and models with the global scientific community. 

Why Ai2:

  • Open Science: Your work directly enables the release of open models like OLMo, providing the broader research community with tools they can't get elsewhere.
  • Mission-Driven: We prioritize scientific impact over profit margins. This allows us to focus on building the "right" infrastructure for long-term research goals.
  • Complexity at Scale: You will manage some of the most dense and high-performance compute environments currently in operation.

Your Next Challenge:

  • Cluster Operations: Manage the availability, performance, and health of our dense on-prem GPU clusters. Coordinate with hardware vendors and internal teams to keep physical infrastructure meeting the demands of frontier model training.
  • Orchestration & Scheduling: Operate and improve Beaker, our internal orchestration platform by optimizing resource allocation and driving high utilization across on-prem assets and elastic cloud resources (AWS/GCP).
  • Storage Operations: Execute and continuously improve our storage environment, balancing high-throughput performance for active training against cost-effective durability for petascale research data. Contribute to the longer-term storage roadmap.
  • Resource Management: Manage GPU compute allocation against budget. Track utilization, surface the data, and recommend when to burst to the cloud versus investing in on-prem capacity, escalating larger trade-offs as needed.
  • User Support & Velocity: Serve as the technical bridge to our research teams. Ensure infrastructure is an accelerator, not a bottleneck, for a diverse set of research objectives.
  • Team Leadership: Manage and grow a team of systems engineers, SREs, and software developers. Set the bar for operational rigor, engineering quality, and a collaborative culture, and keep the team unblocked and delivering.

What You’ll Need:

  • Experience: 12+ years in infrastructure, systems engineering, or HPC (or an advanced degree with 8+ years), including 2+ years supervising a small engineering team (5+).
  • Bachelor's degree in a related field: a relevant advanced degree may substitute for equivalent years of technical work experience.
  • GPU/HPC Stack: Direct experience operating large-scale NVIDIA GPU clusters and high-performance networking (InfiniBand/RoCE).
  • Orchestration: Strong background in Kubernetes, Slurm, or similar orchestration frameworks, particularly in hybrid-cloud configurations.
  • Storage: Hands-on experience with distributed filesystems (e.g., WEKA, Ceph, Lustre) and cloud stora

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

The Allen Institute for AI

View company profile →