Jobs and Careers
CH

Senior Machine Learning Engineer

Chalice Custom Algorithms
Hybrid/RemoteRemotefull_timeVerifiedPosted 22 Dec 2025
💰 $200,000/yr($180,000/yr$200,000/yr)

About the role

About Us

Chalice Custom Algorithms (chalice.ai) is the leading AI application for brands applying their own data and analytics in the real-time decisioning of ad buys. Chalice's software automates data ingestion, predictive analytics, and deployment of custom bidding instructions in a supervised-learning environment, where advertisers can visualize and test their custom algorithms.

Advertisers' algorithms can be deployed across all major DSPs, including The Trade Desk, DV360, as well as Meta and YouTube. Chalice was named "Best Demand Side Tech" by AdExchanger, and powered AdWeek's "Best Use of Programmatic" in the 2023 Media Plan of the Year awards.

We are looking for a highly skilled and experienced Senior Machine Learning Engineer to join our team and play a pivotal role in building our distributed ML infrastructure and advancing our core AI products.

About the Role

We're seeking a Senior Machine Learning Engineer with 5-10 years of industry experience who thrives at the intersection of scalable ML systems, distributed computing, and business impact. In this role, you will develop and deploy production ML models—including neural network architectures for audience modeling and optimization—that directly power our core products: AI Audiences, AI Allocator, CPA Algo, and Curate AI.

You won't be reinventing broken pipelines—you'll be building on a strong foundation designed for scale and maintainability. Our team has already established distributed training infrastructure using Ray + PyTorch on Databricks, with MLflow for experiment tracking and Unity Catalog for data governance. You'll own the lifecycle of ML systems from training and hyperparameter tuning to batch inference and observability.

This is an opportunity to work closely with Directors of Engineering, Product, and Data Science to build systems that directly impact product strategy and business outcomes in the programmatic advertising space.

Key Responsibilities

Distributed Training & Model Development

• Architect, train, and maintain scalable neural network systems for audience modeling and bid optimization using PyTorch and Ray distributed training (Ray Train, Ray Tune, DDP)

• Build and optimize multi-GPU training pipelines on Databricks, including hyperparameter search with ASHA scheduling and early stopping

• Develop feature engineering pipelines using PySpark, including embedding layers (EmbeddingBag, Embedding) for categorical and behavioral features

• Implement model comparison workflows with champion/challenger evaluation on holdout data

MLOps & Production Systems

• Build resilient training and batch inference workflows with a focus on automation, reproducibility, and checkpoint recovery

• Implement robust model monitoring and observability solutions (MLflow, Prometheus, Grafana, Datadog) to track drift, performance metrics (AUC, AUPRC, F1), and system health

• Manage model versioning, experiment tracking, and artifact persistence using MLflow and Unity Catalog

• Work closely with engineering teams to integrate model outputs into production systems and optimize dataflows for fault-tolerance

Technical Leadership & Collaboration

• Partner with product stakeholders to align ML efforts with business impact, KPIs, and product strategy across AI Audiences, AI Allocator, CPA Algo, and Curate AI

• Lead technical design reviews, contribute to internal Python packages, and enforce engineering best practices (testing, CI/CD, modularity)

• Stay current on ML infrastructure advancements (distributed training, inference optimization, model serving patterns) and help guide adoption internally

• Document system architectures, create runbooks, and enable team members to adopt and extend the ML framework

Required Qualifications

• Master's Degree or PhD in Computer Science, Statistics, Machine Learning, or related discipline with 5-10 years of industry experience

Strong proficiency in PyTorch for neural network development, including custom architectures with embedding layers, MLP backbones, and binary classification heads

Production experience with Databricks including Delta Lake, Unity Catalog, Asset Bundles, and cluster management

• Strong grasp of MLOps best practices: experiment tracking (MLflow), model versioning, model serving, monitoring, and reproducibility

• Expert-level Python and PySpark skills for data processing and feature engineering at scale

• Experience building and maintaining batch inference pipelines with schema versioning and artifact management

• Familiarity with cloud platforms (AWS: S3, EC2) and data warehousing (Snowflake)

• Experience with CI/CD workflows

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Chalice Custom Algorithms

View company profile →