Jobs and Careers
MI

Member of Technical Staff, Large Generative Models

Mirage
New York City, United Statesfull_timeVerifiedPosted 22 Oct 2025
💰 $350,000/yr($200,000/yr$350,000/yr)

About the role

Mirage is the leading AI short-form video company. We’re building full-stack foundation models and products that redefine video creation, production and editing. Over 20 million creators and businesses use Mirage’s products to reach their full creative and commercial potential.

We are a rapidly growing team of ambitious, experienced, and devoted engineers, researchers, designers, marketers, and operators based in NYC. As an early member of our team, you’ll have an opportunity to have an outsized impact on our products and our company's culture.

Our Products

Captions

Mirage Studio

Our Technology

AI Research @ Mirage

Mirage Model Announcement

Seeing Voices (white-paper)

Press Coverage

TechCrunch

Lenny’s Podcast

Forbes AI 50

Fast Company

Our Investors

We’re very fortunate to have some the best investors and entrepreneurs backing us, including Index Ventures, Kleiner Perkins, Sequoia Capital, Andreessen Horowitz, Uncommon Projects, Kevin Systrom, Mike Krieger, Lenny Rachitsky, Antoine Martin, Julie Zhuo, Ben Rubin, Jaren Glover, SVAngel, 20VC, Ludlow Ventures, Chapter One, and more.

** Please note that all of our roles will require you to be in-person at our NYC HQ (located in Union Square)

We do not work with third-party recruiting agencies, please do not contact us**

About the role:

Captions is seeking an exceptional Research Engineer (MOTS) to advance the state-of-the-art in large-scale multimodal video diffusion models. You'll conduct novel research on generative modeling architectures, develop new training techniques, and scale models to billions of parameters. As a key member of our ML Research team, you'll work at the cutting edge of multimodal generation while building systems that enable natural, controllable video creation. We're already training large-scale models with demonstrated product impact, and we're excited to continue expanding the scope and capabilities of our research.

We're especially excited about pushing the boundaries of audio-video generation, with a focus on realistic and charismatic human behavior that enables natural storytelling and creative iteration. Our models power creative tools used by millions of creators, and we're tackling fundamental challenges in how to generate compelling human motion, expression, and speech. 

Key Responsibilities:

Research & Architecture Development:

  • Design and implement novel architectures for large-scale video and multimodal diffusion models

  • Develop new approaches to multimodal fusion, temporal modeling, and video control

  • Research temporal video editing techniques and controllable generation

  • Research and validate scaling laws for video generation models

  • Create new loss functions and training objectives for improved generation quality

  • Drive rapid experimentation with model architectures and training strategies

  • Validate research directly through product deployment and user feedback

Model Training & Optimization:

  • Train and optimize models at massive scale (10s-100s of billions of parameters)

  • Develop sophisticated distributed training approaches using FSDP, DeepSpeed, Megatron-LM

  • Design and implement model surgery techniques (pruning, distillation, quantization)

  • Create new approaches to memory optimization and training efficiency

  • Research techniques for improving training stability at scale

  • Conduct systematic empirical studie

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Mirage

View company profile →