Senior Software Development Engineer
ZillowAbout the role
About the team
The Agentic AI team at Zillow is transforming the real estate industry by helping millions of people use AI assistants to find their next home. We are building always-on AI experiences that combine personalized user insights with Zillow’s deep real estate knowledge.The Agentic Evals team focuses on one of the hardest parts of building AI experiences: measuring quality. We build evaluation frameworks, tracing systems, observability tools, and feedback loops that help teams understand what works, identify failure modes, and continuously improve AI assistant behavior. Our work ensures Zillow’s AI agents remain trustworthy, responsible, measurable, and ready to scale.
As part of this lean, customer-focused team, you will partner with applied scientists, software engineers, machine learning engineers, and product leaders to evolve Zillow’s next-generation AI platform.
About the role
We are seeking a collaborative, product-minded backend engineer with strong Agentic AI fundamentals and a passion for building reliable, scalable systems that power AI applications.
In this role, you will:
Design, build, and scale evaluation frameworks for Zillow’s agentic AI experiences.
Build tracing, observability, and quality measurement systems for production AI agents.
Create platform infrastructure that enables domain teams across Zillow to build on top of the Evals platform without bottlenecks.
Partner with applied scientists and machine learning engineers to integrate new AI evaluation capabilities into production systems.
Help evolve how Zillow measures AI quality, reliability, trustworthiness, and product impact.
Stay current with emerging agentic AI paradigms, evaluation techniques, and LLM tooling, and translate them into practical platform innovation.
Support scaling, reliability, performance optimization, incident response, and cost management for the evaluation layer.
Apply first-principles thinking to ambiguous problems and iterate quickly on novel solutions.
Who you are
You are a roll-up-your-sleeves technical expert who can combine state-of-the-art AI technology with large-scale backend engineering. You thrive in ambiguous environments, enjoy solving challenging problems, and care deeply about creating reliable systems that other teams can trust.
You have:
4+ years of backend engineering experience, with a track record of designing, shipping, and operating scalable production ML services
Experience building platform infrastructure or developer-facing services consumed by multiple teams; you've thought carefully about APIs, user experience, reliability, and what it means to have internal customers
Hands-on experience with ML Evals & Observability frameworks (Databricks MLflow, or evolving LLM frameworks preferred - LangSmith, Braintrust, Promptfoo, or similar)
Experience evaluating third-party AI platforms and making principled build vs. buy decisions for platform infrastructure
Fluency working across applied science, ML, product, design, and engineering; you translate between disciplines without losing precision
Comfort with LLMs, agentic systems, and evaluation frameworks;
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s