Jobs and Careers
FI

Member of Technical Staff, DevOps / Infrastructure Engineering

FirstPrinciples
Anywhere - RemoteRemotefull_timeVerifiedPosted 10 Oct 2025

About the role

About FirstPrinciples:
FirstPrinciples is a non-profit organization building an autonomous AI Physicist designed to advance humanity's understanding of the fundamental laws of nature. Our goal for the AI Physicist is to achieve a breakthrough that unifies quantum field theory & general relativity and to explain the deepest unresolved phenomena in our universe by 2035. To do this, we're pioneering a new approach to scientific discovery by creating an intelligent system that can explore theoretical frameworks, reason across disciplines, and generate novel insights. We're a non-profit that operates like a tech start-up by moving quickly and continuously iterating to accelerate scientific progress. By combining AI, symbolic reasoning, and autonomous research capabilities, we're developing a platform that goes beyond analyzing existing knowledge to actively contribute to physics research.

Job Description:
We're seeking a Member of Technical Staff, DevOps / Infrastructure Engineering to architect, automate, and scale the infrastructure that underpins our large-scale model training and research workflows. This role spans both cloud environments (AWS) and HPC infrastructure (Buzz & Lambda HPC GPU clusters with high-speed interconnects), requiring you to design and codify the systems, pipelines, and automation that enable our researchers and engineers to move fast with confidence. Our ideal candidate brings strong fundamentals in Unix/Linux, deep experience in CI/CD and infrastructure-as-code, and a systems mindset to define standards, build automation, and grow our infrastructure practice from the ground up. You'll be instrumental in building the reliable, scalable foundation that powers our autonomous AI Physicist while partnering closely with training engineers and researchers to accelerate breakthrough scientific discoveries.

Key Responsibilities:

Infrastructure Architecture & Automation:

  • Design and run large-scale pre-training experiments for both dense and MoE architectures, from experiment planning through multi-week production runs.
  • Architect hybrid infrastructure solutions that span cloud and on-premises HPC environments seamlessly.
  • Automate configuration management and drift detection using tools like Ansible, Salt, or Chef.
  • Build systems that reduce operational toil and establish guardrails that let researchers focus on experiments, not operations.

CI/CD & Developer Experience:

  • Build and own comprehensive CI/CD pipelines for training workflows, evaluation jobs, internal tools, and services with rollback capabilities, observability, and safety built in.
  • Develop tooling for developer workflows including reproducible builds, ephemeral environments, secrets management, and cluster resource allocation.
  • Create self-service infrastructure patterns that empower researchers and engineers.
  • Design infrastructure that accelerates experimentation while maintaining reliability and reproducibility.

HPC & GPU Cluster Management:

  • Manage and extend HPC environments including GPU clusters, InfiniBand networks, job schedulers (Slurm/Kubernetes hybrid), and container orchestration.
  • Operate containerized and scheduled workloads efficiently across Docker, Kubernetes, and Slurm environments.
  • Optimize cluster scheduling and resource allocation for high-performance GPU workloads.
  • Debug GPU driver quirks, Slurm job issues, and InfiniBand networking hiccups as they arise.

Monitoring, Observability & Reliability: 

  • Implement comprehensive monitoring, logging, and alerting across all infrastructure layers using Prometheus, Grafana, ELK/EFK, and OpenTelemetry.
  • Establish SLOs/SLIs for infrastructure reliability and create observability dashboards for long-horizon training runs.
  • Build observability stacks that provide visibility into both system health and job-level performance.
  • Proactively detect and resolve infrastructure issues before they impact research workflows.

Security & Compliance: 

  • Implement and manage secrets management and identity security solutions (Vault, KMS, IAM).
  • Champion security best practices, IAM policies, and compliance standards across hybrid infrastructure.
  • Design infrastructure with least privilege principles and strong security hygiene from the start.
  • Maintain zero-trust security posture and comprehensive auditing capabilities.

Collaboration: 

  • Partner closely with training engineers and researchers to transla

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

FirstPrinciples

View company profile →