Vice President, AI Platform Engineering
Ares Management CorporationAbout the role
Over the last 20 years, Ares’ success has been driven by our people and our culture. Today, our team is guided by our core values – Collaborative, Responsible, Entrepreneurial, Self-Aware, Trustworthy – and our purpose to be a catalyst for shared prosperity and a better future. Through our recruitment, career development and employee-focused programming, we are committed to fostering a welcoming and inclusive work environment where high-performance talent of diverse backgrounds, experiences, and perspectives can build careers within this exciting and growing industry.
Job Description
Overview
We are seeking an accomplished VP of AI Platform Engineering to lead the design, development, and deployment of our enterprise generative AI platform. This leadership role focuses on building and scaling core platform components that enable safe, secure, and compliant AI application development across the firm. Working closely with the Principal AI Platform Engineer and cross-functional teams, you will drive execution on critical platform infrastructure—from multi-LLM gateways and RAG services to model registry, prompt library, and production deployment pipelines. This is an opportunity to shape how the organization leverages AI at scale while maintaining rigorous standards for security, governance, and reliability.
Key Responsibilities
Platform Development & Execution
- Lead design and implementation of core platform components: multi-LLM gateway, RAG retrieval services, model registry, and prompt library
- Drive execution on platform roadmap, breaking down complex features into deliverable milestones with clear success metrics
- Own API design and service integration patterns that enable seamless consumption across AI enablement teams
- Ensure technical excellence: code quality, testability, performance optimization, and architectural coherence
Multi-LLM Gateway & Model Management
- Design and build multi-LLM gateway architecture supporting multiple providers (OpenAI, Anthropic, Azure, self-hosted, etc.)
- Implement intelligent routing, load balancing, and fallback mechanisms based on cost, latency, and capability requirements
- Build model registry with versioning, metadata management, and approval workflows
- Implement cost optimization and FinOps tracking for model usage and spending
- Monitor model performance, hallucination rates, latency, and quality metrics in production
RAG & Retrieval Infrastructure
- Design and build enterprise RAG infrastructure: vector database integration, semantic search, and chunking strategies
- Implement retrieval evaluation and quality metrics to ensure relevance and accuracy
- Build indexing pipelines and data ingestion workflows from enterprise data sources
- Integrate with data governance and lineage tracking systems
Model Context Protocol (MCP) & Integration Gateway
- Implement MCP gateway for secure, standardized integration with external tools and APIs
- Build tool catalog and discovery mechanisms for AI applications
- Establish security and governance controls for tool access and data handling
Prompt Library & Version Control
- Build organizational prompt library with versioning, tagging, and metadata
- Implement testing and evaluation frameworks for prompt variants
- Enable A/B testing and prompt performance analytics
- Support prompt governance and approval workflows
Deployment Pipelines & DevOps
- Design sandbox-to-production deployment pipelines with clear promotion gates and approval workflows
- Implement CI/CD for AI applications: automated testing, integration, and deployment
- Build monitoring, observability, and alerting for production AI systems
- Implement canary deployments, gradual rollouts, and rollback mechanisms
- Establish SLOs, error budgets, and on-call protocols for platform services
Agent-to-Agent (A2A) Workflows
- Design orchestration framework for multi-step AI workflows with state management
- Build error handling, retries, and recovery mechanisms for reliable execution
- Implement workflow monitoring and debugging tools
Data Integration & Gateway Collaboration
- Partner with Data Products team to design AI-native data access patterns and APIs
- Implement secure, governed data retrieval for RAG and model training
- Build metadata and data lineage tracking for compliance and governance
Security & Governance Implementation
- Implement authentication, authorization, and encryption acr
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s