Jobs and Careers
MI

Senior Data Scientist

Microsoft
United Statesfull_timeVerifiedPosted 30 Oct 2025
💰 $258,000/yr($119,800/yr$258,000/yr)

About the role

M365 Copilot Cadets (Customer & Analytics‑Driven Eval Team) turns real‑world customer feedback into evaluation datasets, rubrics, and insights that measurably improve Microsoft 365 Copilot quality. We connect customer scenarios, analytics, and rigorous evaluation frameworks to power a continuous feedback flywheel across Microsoft 365 Copilot to accelerate measurable product improvements.


As a Senior Data Scientist part of Cadets, you will own evaluation analytics end‑to‑end: curate datasets from customer and production signals; author binary‑first rubrics; build LLM (Large Language Model)‑as‑judge graders and work on high‑quality synthetic data generation to scale evaluations with experience in human‑match rates. You’ll partner with PM/Eng/Design and VIP customers to ship quality gains and AI features with confidence.

You’ll Thrive Here If You Have:Evaluation proficiency for LLM/agent systems: dataset curation, rubric design, human‑in‑the‑loop grading, judge prompts with quantitative agreement goals.

Experience in analytics & experimentation skills (statistical inference, A/B), plus Python/SQL for large‑scale trace analysis.

LLM fundamentals: prompt engineering, few‑shot design, retrieval metrics, multi‑turn/agent trace evaluation.

Data quality mindset: trace hygiene, metadata design, policy/PII awareness, and principled guardrails.

 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

  • Evaluation & Feedback Analysis
  • Convert multi‑source feedback (dogfood, VIP customers, production traces) into a prioritized dataset of 10–100 tasks per scenario, each with prompts and golden outputs; maintain a living failure taxonomy prioritized by volume × impact × fixability.
  • Rubrics & LLM‑as‑Judge
  • Author crisp, binary‑first rubrics across 7–30 dimensions (e.g., correctness/completeness, refusal calibration, tool‑use quality, formatting/contract, persona/tone, trace hygiene).
  • Build grader prompts (with few‑shots and counter‑examples) that achieve ≥80% human‑match rate, track TPR/TNR on held‑out sets, and prevent reward hacking.
  • Synthetic & Human‑Labeled Data
  • Design structured tuples to scale high‑signal synthetic data; orchestrate vendor/partner annotation sprints and live calibrations to align shared judgment.
  • Ensure datasets are reproducible with linked artifacts and robust metadata/trace hygiene.
  • Customer‑Grounded Scenarios
  • Partner with PMs/solution architects to co‑develop evals with VIP customers so tasks reflect real outcomes and workflows; quantify lift from fixes and inform the next hill‑climb.
  • Team Leadership & Ways of Working
  • Co‑own the Cadets “feedback flywheel” with PM/Eng (instrumentation, taxonomy, guardrails vs. evaluators) and help operationalize weekly checklists, change logs, and judge refresh cadence.

Qualifications

Required Qualifications: 

  • Doctorate in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 1+ year(s) data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results)
    • OR Master's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 3+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results)
    • OR Bachelor's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 5+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results)
    • OR equivalent experience.
  • Experience with building data pipelines, performing large-scale analysis, and implementing ML workflows using Python and SQL.
  • Experience in developing models or designing evaluation frameworks, including A/B testing or prompt-based assessments for LLMs.

Other Requirements:
Abi

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Microsoft

View company profile →