AI Trust Data Scientist II
Spring HealthAbout the role
Our mission: to eliminate every barrier to mental health.
At Spring Health, we’re on a mission to revolutionize mental healthcare by removing every barrier that prevents people from getting the help they need, when they need it. Our clinically validated technology, Precision Mental Healthcare, empowers us to deliver the right care at the right time—whether it’s therapy, coaching, medication, or beyond—tailored to each individual’s needs.
We proudly partner with over 450 companies, from startups to multinational Fortune 500 corporations, as a leading provider of mental health service, providing care for 10 million people. Our clients include brands you use and know like Microsoft, Target, and Delta Airlines, all of whom trust us to deliver best-in-class outcomes for their employees globally. With our innovative platform, we’ve been able to generate a net positive ROI for employers and we are the only company in our category to earn external validation of net savings for customers.
We have raised capital from prominent investors including Generation Investment, Kinnevik, Tiger Global, Northzone, RRE Ventures, and many more. Thanks to their partnership and our latest Series E Funding, our current valuation has reached $3.3 billion. We’re just getting started—join us on our journey to make mental healthcare accessible to everyone, everywhere.
As a Data Scientist II on the AI Trust team, you will be a key driver in ensuring our artificial intelligence systems are safe, reliable, and effective. You will own and conduct critical analyses and experiments to measure the real-world impact of our AI, helping to define and validate that our systems are technically robust, trustworthy, and outcome-driven. In this deeply cross-functional position, you will collaborate with partners in engineering, product, and other functions to implement the standards and systems that make "safe and effective" AI a reality.
Please note that candidates for this position must be based in the Salt Lake City metro area and be willing to commute 2-3 days a week when this role transitions to a hybrid schedule in 2026. We're excited to be growing our presence in Salt Lake City!
What you’ll do:
- Own and evolve the evaluation frameworks for our AI and ML models, translating high-level trust principles into specific, measurable tests.
- Define and conduct rigorous experiments to resolve ambiguous questions about the safety, reliability, and impact of our models.
- Collaborate with engineering partners to design and build production-quality code, creating automated, scalable, and pragmatic testing frameworks based on modern best practices.
- Partner with product, legal, and infrastructure teams to implement and monitor standards for trustworthy AI.
- Proactively identify gaps and develop novel evaluation approaches, which may include creating synthetic test data from user traces or building lightweight processes for non-technical partners to iterate on test sets.
- Synthesize complex evaluation results and industry trends into actionable insights and clearly communicate findings to diverse technical and non-technical stakeholders.
What success looks like:
- Measurable improvements in key performance and safety metrics on a quarterly basis.
- Successful design and delivery on POC experiments to improve our base LLM systems.
- Timely and successful delivery of AI evaluation readouts and safety reviews that directly inform key product improvements
- Demonstrated impact on our AI strategy through proactive identification of risks and opportunities for enhancement, confirmed by stakeholder feedback.
- Increased efficiency and coverage of our AI evaluation process, driven by your ownership of and improvements to automated testing frameworks and best practices.
What you’ll bring:
- 2-3 + years of relevant industry experience in data science, machine learning, or a related field.
- Proficiency in Python and a solid understanding of core statistical concepts. You have a proven ability to write and review production-quality code.
- Proven experience in evaluating machine learning models, with exposure to large language models (LLMs) being a strong plus.
- Hands-on experience in one or more of the following areas:
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s