Senior Staff ML Engineer
ZendeskAbout the role
Job Description
Zendesk’s people have one goal in mind: to make Customer Experience better. Our products help more than 125,000 global brands (AirBnb, Uber, JetBrains, Slack, among others) make their billions of customers happy, every day.
The AI/ML Platform team is at the forefront of this mission. We build the foundation that powers every AI-driven experience at Zendesk, enabling product teams to build, evaluate, and deploy state-of-the-art Large Language Model (LLM) applications reliably and at scale.
We’re looking for a Senior Staff ML Engineer to set the technical vision and architecture for Zendesk’s next generation of GenAI infrastructure and platform. You’ll be the primary technical authority for systems like our LLM Proxy, model evaluation frameworks, agent orchestration tools, and multi-model routing ensuring they are secure, scalable, performant, cost-efficient, and future-proof.
This is a hands-on architecture and leadership role, you’ll design high-impact systems, mentor senior engineers, and work across product and platform teams to ensure GenAI capabilities are consistently delivered with excellence.
What you get to do every day
Own the end-to-end architecture for Zendesk’s GenAI platform, ensuring alignment with business goals and technical best practices.
Set technical direction for core systems including LLM Proxy, agent orchestration layers, evaluation and benchmarking pipelines, and model observability tooling
Design and implement model routing, fallback strategies, and A/B testing infrastructure for LLMs from multiple vendors
Establish engineering standards for safety, latency, cost attribution, and reliability across all GenAI services.
Coach Staff and Senior Engineers, providing deep technical guidance and fostering a culture of technical excellence.
What you bring to the role
10+ years in ML/AI engineering, including at least 3 years in staff/principal-level platform or infrastructure leadership roles..
Expertise in LLM systems, multi-model orchestration, and GenAI infrastructure patterns
Proven ability to design complex, distributed systems with high reliability, security, and scalability requirements
Strong experience with AWS, GCP, or Azure; Kubernetes; Docker; and distributed event-driven architectures
Fluency in Python and at least one other server-side language (Java, Scala, Golang, or Ruby).
Track record of driving alignment across multiple engineering teams and influencing product direction.
Preferred Qualifications
Background in agentic architectures and complex workflow orchestration for AI agents.
Experience and contributions to enterprise-scale ML platforms.
What our tech stack looks like
Our code is written in Python.
Our servers live in AWS.
LLM Vendors: OpenAI, Anthropic, Google, Llama
Infra: Kubernetes, Docker, Kafka, AWS
What we offer
Full ownership of the projects you work on.
What you will be doing will have a huge impact.
Team of passionate people who love what they do.
Exciting projects, ability to implement your own ideas and improvements.
Opportunity to learn and grow.
...and everything you need to be effective and maintain work-life balance
Flexible working hours.
Professional development funds.
Comfortable office and a remote setup.
Choice of your laptop and other equipment.
Premium Medical Insurance as well as Private Life Assurance.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s