Staff ML Engineer
Group 1001About the role
Group 1001 is a consumer-centric, technology-driven family of insurance companies on a mission to deliver outstanding value and operational performance by combining financial strength and stability with deep insurance expertise and a can-do culture. Group1001’s culture emphasizes the importance of collaboration, communication, core business focus, risk management, and striving for outcomes. This goal extends to how we hire and onboard our most valuable assets – our employees.
Why This Role Matters:
We're building AI&ML-powered products that will transform how Group 1001 approaches pricing optimization, claims automation, and risk intelligence. To do this at scale, we need robust ML infrastructure—not just great models.
As a Staff ML Engineer, you'll focus on the MLOps and infrastructure layer that makes ML production-ready: model serving, feature pipelines, experiment tracking, and CI/CD for ML. You'll help shape our ML platform architecture, working alongside Platform Engineering teams to ensure ML workloads run reliably on our modern stack: Snowflake, Dagster, Coalesce, Palantir and AWS SageMaker.
This role is for engineers who are as passionate about infrastructure, deployment, and operationalizing ML as they are about the models themselves
*Please note, this position requires an in-person interview.
How You'll Contribute:
- Partner with Data & Platform Engineering to define how ML workloads integrate with our Snowflake-Dagster-Palantir ecosystem
- Evaluate and recommend tooling for the ML stack—balancing build vs. buy decisions against our scale and compliance needs
- Contribute to platform roadmap discussions, advocating for infrastructure investments that accelerate ML delivery
- Establish CI/CD pipelines for ML: automated testing, model validation, staged deployments, and rollback capabilities using SageMaker Pipelines, Step Functions, or similar orchestration
- Implement model monitoring and observability: drift detection, performance degradation alerts, and automated retraining triggers
- Architect ML workloads on AWS: SageMaker (Training Jobs, Processing, Endpoints), EC2/EKS for custom serving, S3 for artifact storage, and IAM for secure access patterns
- Optimize for cost and performance—right-sizing instances, spot instance strategies, auto-scaling endpoints, and efficient GPU utilization
- Integrate ML infrastructure with our Dagster orchestration layer for end-to-end pipeline visibility
- Mentor senior ML engineers and technical leads, developing the next generation of ML engineering leadership
What We're Looking For:
Technical Skills:
- MLOps & Model Serving: Hands-on experience with model serving frameworks (SageMaker Endpoints, Seldon Core, BentoML, Ray Serve, or TensorFlow Serving); building and operating inference infrastructure at scale
- CI/CD for ML: Building ML pipelines with SageMaker Pipelines, Kubeflow, Airflow, or Dagster; automated model testing, validation gates, and deployment automation
- AWS & Cloud Infrastructure: Strong AWS experience—SageMaker, EKS/ECS, Lambda, Step Functions, S3, IAM; infrastructure-as-code (Terraform, CDK, CloudFormation)
- Monitoring & Observability: Model monitoring, drift detection, alerting; tools like Evidently, WhyLabs, SageMaker Model Monitor, or custom solutions
- Core ML Fundamentals: Working knowledge of Python, ML frameworks (PyTorch, TensorFlow, scikit-learn), and model evaluation—enough to partner effectively with data scientists
- Feature Engineering Infrastructure: Experience with feature stores (SageMaker Feature Store, Feast, Tecton, or similar); designing feature pipelines for both batch and real-time serving
- Experiment Tracking & Registry: MLflow, Weights & Biases, SageMaker Experiments, or similar; establishing reproducibility and governance across ML projects
- Nice to Have: Palantir Foundry, Kubernetes, Bedrock, cost optimization strategies for ML workloads
Education:
- Bachelor's degree in Computer Science, Data Science, Engineering, or related field
- Master's degree or equivalent experience preferred
Experience:
- 7-10 years in ML engineering, MLOps, or platform engineering with a focus on productionizing ML systems
- Demonstrated experience building ML infrastructure that others build upon—serving layers, feature stores, or MLOps tooling
- Track record of improving ML delivery velocity through infrastructure and automation
- Proven ability to work cross-functionally with data scientists, platform engineers, and stakeholders
- Experience mentoring and developing senio
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s