Machine Learning Engineering, Training Data Infrastructure
CaptionsAbout the role
Captions is the leading video AI company, building the future of video creation. Over 10 million creators and businesses have used Captions to create videos for social media, marketing, sales, and more. We're on a mission to serve the next billion.
We are a rapidly growing team of ambitious, experienced, and devoted engineers, researchers, designers, marketers, and operators based in NYC. You'll join an early team and have an outsized impact on the product and the company's culture.
Weβre very fortunate to have some the best investors and entrepreneurs backing us, including Index Ventures (Series C lead), Kleiner Perkins (Series B lead), Sequoia Capital (Series A and Seed co-lead), Andreessen Horowitz (Series A and Seed co-lead), Uncommon Projects, Kevin Systrom, Mike Krieger, Lenny Rachitsky, Antoine Martin, Julie Zhuo, Ben Rubin, Jaren Glover, SVAngel, 20VC, Ludlow Ventures, Chapter One, and more.
Check out our latest financing milestone and some other coverage:
The Information: 50 Most Promising Startups
Fast Company: Next Big Things in Tech
The New York Times: When A.I. Bridged a Language Gap, They Fell in Love
Business Insider: 34 most promising AI startups
Time: The Best Inventions of 2024
** Please note that all of our roles will require you to be in-person at our NYC HQ (located in Union Square) **
Overview
Captions seeks an exceptional Machine Learning Engineer to drive innovation in training data infrastructure. You'll conduct research on and develop sophisticated distributed training workflows and optimized data processing systems for massive video and multimodal datasets. Beyond pure performance, you'll develop deep insight into our data to maximize training effectiveness. As an early member of our ML Research team, you'll build foundational systems that directly impact our ability to train models powering video and multimodal creation for millions of users.
Key Responsibilities
Infrastructure Development:
Build performant pipelines for processing video and multimodal training data at scale
Design distributed systems that scale seamlessly with our rapidly growing video and multimodal datasets
Create efficient data loading systems optimized for GPU training throughput
Implement comprehensive telemetry for video processing and training pipelines
Core Systems Development:
Create foundation data processing systems that intelligently cache and reuse expensive computations across the training pipeline
Build robust data validation and quality measurement systems for video and multimodal content
Design systems for data versioning and reproducing complex multimodal training runs
Develop efficient storage and compute patterns for high-dimensional data and learned representations
System Optimization:
Own and improve end-to-end training pipeline performance
Build systems for efficient storage and retrieval of video training data
Build frameworks for systematic data and model quality improvement
Develop infrastructure supporting fast research iteration cycles
Build tools and systems for deep understanding of our training data characteristics
Research & Product Impact:
Build infrastructure enabling rapid testing of research hypotheses
Create systems for incorporating user feedback into training workflows
Design measurement frameworks that connect model improvements to user outcomes
Enable systematic experimentation with direct user feedback loops
Preferred Qualifications:
Technical Background:
Bachelor's or Master's degree in Computer Science, Machine Learning, or related field
3+ years experience in ML in
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights β in under 60 seconds.
Apply Now βGenerate Application KitFree account required β sign up in 30s