Senior Engineer, Inference Control Plane
DigitalOceanAbout the role
<div class="content-intro"><p>Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world. </p></div><p>We are seeking a Senior Engineer to implement and contribute to the design and optimization of our Serverless Inference infrastructure and APIs. In this role, you will tackle the challenges of large-scale AI workloads, focusing on throughput, GPU utilization, and fault tolerance to support next-generation inference needs of AI native enterprises.</p> <h2><strong>What You'll Do:</strong></h2> <ul> <li>Design and build scalable, multi-tenant services that power AI inference and intelligent routing workloads.</li> <li>Develop and operate high-scale distributed systems with strong reliability, availability, and performance goals.</li> <li>Strengthen platform resiliency through improved observability, capacity management, automation, and operational tooling.</li> <li>Partner closely with platform, GPU infrastructure, and product engineering teams to deliver production-grade systems and highly available APIs.</li> <li>Raise the engineering bar through strong software design, operational discipline, incident management, and continuous improvement practices.</li> <li>Contribute to architecture decisions around traffic management, service orchestration, reliability, and platform scalability.</li> <li>Participate in on-call rotations and lead efforts to reduce operator pain, improve service health, and prevent recurring incidents.</li> </ul> <h2><strong>What You'll Bring:</strong></h2> <p><strong>Required </strong></p> <ul> <li>5+ years of experience building and operating multi-tenant platforms or distributed backend systems</li> <li>Strong experience operating high-scale distributed services in production environments</li> <li>Deep understanding of SRE principles, including observability, incident management, reliability engineering, capacity planning, and operational automation</li> <li>1+ years of hands-on experience with Go / Golang in production systems</li> <li>1+ years of experience with Kubernetes</li> <li>Strong understanding of cloud-native architectures, microservices, and distributed systems fundamentals</li> <li>Experience debugging performance, scalability, and reliability issues in producti
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s