Jobs and Careers
FI

Member of Technical Staff, Training Engineer (Large Scale Foundation Models)

FirstPrinciples
Anywhere - RemoteRemotefull_timeVerifiedPosted 10 Oct 2025

About the role

About FirstPrinciples:
FirstPrinciples is a non-profit organization building an autonomous AI Physicist designed to advance humanity's understanding of the fundamental laws of nature. Our goal for the AI Physicist is to achieve a breakthrough that unifies quantum field theory & general relativity and to explain the deepest unresolved phenomena in our universe by 2035. To do this, we're pioneering a new approach to scientific discovery by creating an intelligent system that can explore theoretical frameworks, reason across disciplines, and generate novel insights. We're a non-profit that operates like a tech start-up by moving quickly and continuously iterating to accelerate scientific progress. By combining AI, symbolic reasoning, and autonomous research capabilities, we're developing a platform that goes beyond analyzing existing knowledge to actively contribute to physics research.

Job Description:
We're seeking a Member of Technical Staff, Training Engineer to develop and lead end-to-end pre-training of large language models on GPU clusters as we build the AI Physicist to revolutionize fundamental physics research. You'll make critical modeling choices, guide the development of data pipelines, and perform distributed training at scale, all guided by rigorous evaluation frameworks. This role requires you to combine deep engineering expertise with research intuition to push throughput, stability, and final capability while productionizing successful ideas into repeatable training runs and reusable tooling. You'll be instrumental in building the foundation models that power the AI Physicist, ensuring every training run brings us closer to breakthrough scientific discoveries.

Key Responsibilities:

Model Training & Optimization:

  • Design and run large-scale pre-training experiments for both dense and MoE architectures, from experiment planning through multi-week production runs.
  • Tune optimizer configurations (AdamW/Adafactor/Sophia variants), learning rate schedules with warmup strategies, dropout, gradient clipping, weight decay, EMA, and activation checkpointing to ensure stability at scale.
  • Own model and training recipes end-to-end, making informed decisions about microbatch and global batch configurations.
  • Run ablations and scaling-law studies to set optimal tokens-to-train targets, entropy/perplexity goals, and checkpoint cadence that optimize cost-to-quality ratios.
  • Provide strategic insights to the executive team on financial implications of major decisions, from international expansion to new research initiatives.
  • Design capital allocation frameworks that maximize scientific impact while ensuring long-term sustainability.

Data Pipeline Engineering:

  • Build and harden high-throughput data pipelines encompassing dataset curation, filtering, deduplication, pack-by-length optimization, and contamination control.
  • Design and implement multilingual and multimodal data ingest systems with intelligent repeat scheduling (e.g., D4-style approaches).
  • Architect comprehensive data pipelines across diverse modalities (web/book/code/speech/vision) with filtering, heuristic and learned scoring, temperature sampling, multilingual balancing, and curriculum learning.
  • Demonstrate measurable impact from data quality work including large-scale deduplication, contamination audits, and repeat/mixture scheduling that improves downstream accuracy.

Distributed Training & Performance:

  • Operate distributed training infrastructure using FSDP/ZeRO, tensor/pipeline/expert/context parallelism, and high-speed interconnects (NCCL, NVLink/InfiniBand).
  • Choose and configure optimal distributed strategies (FSDP vs ZeRO; 3D/5D hybrid parallelism for MoE) and launch parameters, documenting trade-offs for future reference.
  • Exploit modern kernels and mixed-precision training (FlashAttention-3, FP8 via NVIDIA Transformer Engine) to maximize tokens/sec while maintaining perplexity targets.
  • Integrate performance primitives including FlashAttention-3, fused optimizers, and custom CUDA/Triton kernels while maintaining convergence guarantees.
  • Write production-grade PyTorch and Triton/CUDA kernels when required to unlock critical performance gains.

Reliability & Observability: 

  • Debug complex distributed training issues including deadlocks, OOMs, divergence, and stragglers using tools like Nsight, py-spy, TensorBoard, and W&B.
  • Build comprehensive observability systems for long-horizon runs tracking throughput/efficiency, gradient statistics, loss spikes, token-mix drift, dat

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

FirstPrinciples

View company profile →