Member of Technical Staff — Supercomputing
RadixArkAbout the role
<h2><strong>About the Role</strong></h2> <p>RadixArk is hiring a <strong>Member of Technical Staff — Supercomputing</strong> to help build, deploy, and operate production-grade AI infrastructure for frontier-scale inference and training workloads.</p> <p>This role sits at the intersection of engineering, deployment, reliability, and customer infrastructure. You will work on bringing up SGLang, Miles, and the RadixArk infrastructure stack across cloud GPUs, customer VPCs, dedicated clusters, and partner environments. You will help ensure that our systems are not only fast in benchmarks, but reliable, observable, and operationally robust in real production settings.</p> <p>This is not a traditional DevOps or SRE role. You will need to understand the model, the serving engine, the cluster, the workload, and the production constraints. One day you may be helping a customer bring up a new model on a GPU cluster; another day you may be debugging P99 latency, GPU utilization, autoscaling, networking behavior, deployment failures, or reliability regressions.</p> <p>We are looking for someone hands-on, technically strong, calm under pressure, and excited to build the supercomputing foundation for a new category of AI infrastructure company.</p> <h2><strong>What You’ll Do</strong></h2> <ul> <li>Deploy SGLang, Miles, and RadixArk infrastructure across customer, cloud, VPC, and dedicated cluster environments.</li> <li>Bring up production inference and training workloads for open-weight and customer-specific models.</li> <li>Own deployment reliability, environment management, rollout processes, and production validation.</li> <li>Debug issues across LLM serving, Kubernetes, networking, GPU infrastructure, storage, cloud capacity, and customer systems.</li> <li>Build and improve observability for latency, throughput, uptime, error rates, GPU utilization, memory usage, capacity, and workload health.</li> <li>Improve monitoring, alerting, incident response, runbooks, postmortems, and operational processes.</li> <li>Help design capacity planning, autoscaling, and reliability strategies for GPU-intensive workloads.</li> <li>Work closely with engineering teams to improve deployment tooling, automation, CI/CD, and production readiness.</li> <li>Partner with customer engineering teams during POCs, production launches, and ongoing operations.</li> <li>Feed deployment and reliability pain points back into the product and engineering roadmap.</li> <li>Help build the foundation for a world-class supercomputing deployment and reliability organization.</li> </ul> <h2><strong>What We’re Looking For</strong></h2> <ul> <li>Strong hands-on experience operating production systems at meaningful s
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s