Jobs and Careers
AP
Senior Software Engineer
AppleUnited Statesfull_timeVerifiedPosted 23 Jul 2026
About the role
Summary
As part of the Siri organization, you will build the systems and tooling that make evaluation a first-class part of how Siri is developed — not an after-the-fact check — spanning human evaluation, real user feedback, reward and alignment signals, and data science rigor across iOS, iPadOS, macOS, watchOS, and visionOS.This is a rare opportunity to work at the intersection of software engineering and rigorous evaluation science — applying machine learning engineering techniques, from model evaluation to reward modeling, to build the infrastructure that keeps this quality signal trustworthy. What you build will directly shape the direction of one of the world's most widely used assistants.
Description
In this role you'll contribute across several interconnected work streams spanning evaluation quality, reward/alignment signals, and data science. A core part of the job is bringing "evals first" thinking to the team — building tooling and harnesses grounded in real workflows, and designing for observability and reproducibility so quality can be measured clearly and issues caught early. This is a largely unexplored space with few established playbooks, so being self-driven is a must — you'll define your own path as much as execute one. Scope and priorities will evolve, and we're looking for someone who moves fluidly across these areas, bringing strong software engineering fundamentals with enough ML/LLM depth to build and ship AI-facing toolingPreferred Qualifications
Experience evaluating ML, LLM, or agent-based systems, including familiarity with metrics, scoring methodology, trajectory and outcome analysis, and techniques like prompting, RAG, or LLM as judgeUnderstanding of reinforcement learning and the underlying techniques behind modern LLMs (e.g. transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design
Familiarity with eval-driven development — defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks
Experience with data science methods applied to quality measurement — defining ground truth, measuring inter-rater agreement (e.g. Cohen's/Fleiss' kappa), and validating automated scorers using basic statistical techniques (e.g. correlation, confidence intervals, hypothesis testing)
Experience with MLOps, deployment, and test/eval environment management — containerization, CI/CD, model versioning, monitoring, cloud platforms (AWS, GCP, or similar), and staging or provisioning environments to produce repeatable, deterministic conditions
Comfort communicating and collaborating effectively across multicultural teams and time zones
Minimum Qualifications
Strong programming skills in one or more compiled languages (Swift, C++, or Objective-C)Strong Python skills and solid computer science fundamentals, including data structures, algorithms, and clean, testable code
Ability to quickly learn and adapt to evolving technologies and tools, such as GenAI-assisted coding, new ML frameworks, and emerging LLM/agent tooling
Experience with backend/API development and production debugging
Excellent communication and cross-team collaboration skills, with experience working effectively within large, cross-functional organizations
M.S. or B.S. in Computer Science, Machine Learning, or a related field (or equivalent experience)
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s