Staff Engineer - Machine Learning
FreshworksAbout the role
Company Description
Organizations everywhere struggle under the crushing costs and complexities of “solutions” that promise to simplify their lives. To create a better experience for their customers and employees. To help them grow. Software is a choice that can make or break a business. Create better or worse experiences. Propel or throttle growth. Business software has become a blocker instead of ways to get work done.
There’s another option. Freshworks. With a fresh vision for how the world works.
Freshworks Inc. builds uncomplicated service software that delivers exceptional employee and customer experiences. Our people-first approach to AI eliminates friction, helping businesses reduce complexity, lower cost-to-serve, and deliver faster, more human support through enterprise-grade yet easy-to-use CX and IT solutions. Nearly 75,000 companies, including Bridgestone, New Balance, Nucor, S&P Global, and Sony Music, trust Freshworks to power their Employee Experience (EX) and Customer Experience (CX) operations.
Fresh vision. Real impact. Come build it with us.
Job Description
We are seeking a Machine Learning Staff Engineer to lead the development of core backend services powering our Agentic AI Platform. In this role, you will be the primary technical driver for building reasoning-driven agents, multi-agent orchestration, and outcome-based workflows.
You will collaborate with the Principal AI Architect to bridge the gap between high-level design and scalable implementation. You will own the development of high-performance APIs and orchestration runtimes, integrating frameworks like LangChain, LangGraph, and LangSmith. This role is highly hands-on and requires a technical leader who can navigate a fast-paced environment while maintaining high engineering standards.
Key Responsibilities
Platform Backend Development
- Lead Implementation: Act as the lead engineer for agent runtime orchestration services, focusing on reasoning, planning, and tool invocation
- API & SDK Design: Own the development of robust APIs and SDKs that enable internal and external teams to build and deploy agent workflows
- State & Memory Management: Build and optimize stateful dialog management and memory services for multi-turn, context-aware agents
- Agent Communication: Implement multi-agent communication (A2A) protocols and shared context systems
System Architecture & Execution
- Microservices Leadership: Execute the transition to cloud-native, event-driven backend services using microservices or service mesh architectures
- Workflow Design: Build and manage task scheduling and long-running operations using tools like Temporal or Airflow
- Performance Optimization: Optimize LLM orchestration for performance and cost through caching, batching, and token monitoring
Integrations & Data Services
- Data Pipeline Ownership: Implement and maintain RAG pipelines and integrations with vector databases (e.g., Pinecone, Weaviate, FAISS)
- Telemetry Integration: Ensure all backend services support LangSmith-based evaluation and comprehensive observability
Security & Reliability
- Tenant Isolation: Implement RBAC, API authentication, and multi-tenant isolation logic
- Engineering Excellence: Set the standard for automated testing (unit, integration, load) and implement OpenTelemetry for full system visibility
- Mentorship: Provide technical guidance and code reviews for more junior backend engineers on the team
Qualifications
Required Qualifications:
- 6+ years of backend engineering experience
- 1-2+ years of hands-on experience in AI/ML platform development
- Expert Proficiency: Strong programming skills in Java and Python
- Framework Expertise: Hands-on experience with LangChain, LangGraph, or LangSmith
- Distributed Systems: Deep understanding of event-driven architectures, microservices, and Kubernetes
- Cloud & Data: Solid experience with AWS, Terraform, and a mix of relational (PostgreSQL) and Vector databases
- Orchestration: Practical experience with workflow engines like Temporal or Airflow and messaging systems like Kafka
Preferred Qualifications:
- Experience with AI-driven planning systems or dialog state tracking
- Background in enterprise security (RBAC, SSO, OAuth2)
- Experience working in a high-growth SaaS or startup environment
- Active contributions to open-source backend or AI frameworks
Additional Information
Please note this is a hybrid role with onsite expectations of 3x/week (Tues - Thurs) from our San Mateo, CA headquarters.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s