Senior Manager - AI Infrastructure & GPU Engineering
EYAbout the role
Location: Alpharetta, Atlanta
At EY, we’re all in to shape your future with confidence.
We’ll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world.
Senior Manager – AI Infrastructure & GPU Engineering
The opportunity
We’re seeking a seasoned Senior Manager to architect and lead the development of GPU-based infrastructure for scalable AI workloads. This role is ideal for a technical leader with deep expertise in infrastructure and DevOps, who thrives in high-performance environments and is passionate about enabling reliable, production-grade AI systems.
Your key responsibilities
- Architect and manage GPU infrastructure to support scalable AI workloads.
- Lead implementation of NVIDIA enterprise software including CUDA, Triton Inference Server, and NIM/NCP services.
- Provide DevOps and AIOps leadership across CI/CD, observability, and reliability engineering practices.
- Mentor infrastructure-focused Managers on system performance optimization and best practices.
Skills and attributes for success
- Strategic thinking with hands-on technical leadership.
- Strong troubleshooting and performance tuning capabilities.
- Ability to lead cross-functional teams and drive infrastructure excellence.
- Passion for innovation and continuous improvement in AI systems.
- Embrace emerging technologies and digital tools to enhance collaboration, streamline workflows, and improve service delivery.
- Promote a culture of continuous learning and adaptability in a hybrid work environment.
- Lead with insight and analytical vigor.
- Apply complex problem-solving and critical thinking to evaluate data, identify trends, and make informed decisions.
- Translate insights into actionable strategies that drive measurable outcomes.
To qualify you must have
- 8–12+ years of experience in infrastructure or platform engineering.
- Strong knowledge of the NVIDIA GPU ecosystem, including MIG, Run.ai, and Kubernetes operators.
- Experience with Terraform, Ansible, Helm, and monitoring stacks such as Prometheus and Grafana.
- Familiarity with hybrid cloud environments (Azure, AWS, GCP) for GPU workloads.
Ideally, you'll also have
- Experience deploying and scaling AI infrastructure in enterprise settings.
- Knowledge of security and compliance considerations for AI platforms.
- Excellent communication and stakeholder engagement skills.
What we offer you
At EY, we’ll develop you with future-focused skills and equip you with world-class experiences. We’ll empower you in a flexible environment, and fuel you and your extraordinary talents in a diverse and inclusive culture of globally connected teams. Learn more.
- We offer a comprehensive compensation and benefits package where you’ll be rewarded based on your performance and recognized for the value you bring to the business. The base salary range for this job in all geographic locations in the US is $144,000 to $329,100. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $172,800 to $374,000. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography. In addition, our Total Rewards package includes medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options.
- Join us in our team-led and leader-enabled hybrid model. Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s