Senior Manager, AI Agent / ML Engineering
OracleAbout the role
The Senior Manager will lead a team of engineers while serving as a hands-on technical leader responsible for defining, building, and operating next-generation AI systems on Oracle Cloud Infrastructure (OCI). This leader will set the architecture and engineering direction for production-grade agentic AI platforms, autonomous workflows, scalable inference infrastructure, and enterprise AI applications used in large-scale, business-critical environments.
This role requires an experienced engineering manager who can build and develop high-performing teams, translate ambiguous product and platform goals into a durable technical strategy, and drive execution across multiple organizations. The successful candidate will be accountable for hiring, mentoring, performance management, technical planning, and delivery, while remaining actively involved in system design, prototyping, coding, code reviews, operational readiness, and incident follow-up.
The ideal candidate combines deep distributed-systems expertise with practical, hands-on experience building AI agents and AI-native applications. This includes developing and orchestrating LLM-based agents, tools, APIs, memory systems, retrieval pipelines, evaluations, guardrails, and cloud-service integrations. The candidate should be comfortable writing production code, debugging complex systems, and guiding engineers through difficult architectural and implementation decisions.
The expectation is to lead the team in shipping, scaling, and operating reliable, secure, observable, and cost-efficient AI systems, while raising both the engineering and management bar across the organization. This leader will establish strong execution practices, promote operational excellence, and ensure the team delivers measurable business and customer outcomes.
Responsibilities
- Lead and develop a team responsible for OCI AI platform capabilities, including agent execution, inference, orchestration, evaluation, and observability.
- Set the technical direction for production-grade agentic AI systems that support reasoning, planning, tool use, multi-step workflows, and human escalation.
- Remain hands-on in architecture, prototyping, coding, debugging, and code reviews for critical AI-agent components.
- Guide the development of services for tool calling, memory, context management, MCP integration, retrieval, multi-agent coordination, policy enforcement, and evaluation.
- Own delivery across distributed systems optimized for reliability, performance, security, cost, and multi-tenant operation.
- Translate broad goals into roadmaps, staffing plans, milestones, and measurable outcomes.
- Partner across infrastructure, security, data, product, and application teams to drive execution.
- Establish AgentOps and LLMOps practices for tracing, monitoring, testing, safety guardrails, versioning, and production readiness.
- Recruit, coach, and retain engineers while managing performance and developing senior technical leaders.
- Own production outcomes, including reliability, security, cost efficiency, supportability, and delivery predictability.
Required Qualifications
- Bachelor's, Master's, or Ph.D. in Computer Science, AI/ML, Engineering, or a related field, or equivalent experience.
- 8+ years of software engineering experience, including ownership of production systems.
- 2+ years of engineering management experience, including hiring, coaching, performance management, and delivery ownership.
- Proven ability to lead teams while remaining technically engaged in design, coding, reviews, debugging, and operations.
- Deep experience with distributed systems, cloud platforms, or AI/ML infrastructure.
- Hands-on experience building AI agents, autonomous workflows, tool-using systems, or multi-step orchestration.
- Experience with frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar tools.
- Strong understanding of LLM patterns, including tool calling, RAG, memory, context management, evaluation, and safety.
- Strong Python skills and experience with Kubernetes, Docker, observability, scalability, and fault tolerance.
- Strong understanding of AI security, governance, access control, auditability, and operational risk.
- Excellent communication and cross-functional leadership skills.
Preferred Qualifications
- Experience managing teams that build AI platforms, agent runtimes, inference systems, or devel
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s