Senior AI Workload Platform Engineer - Radian Arc
SubmerAbout the role
Location & work modality: EMEA (remote)
Start: ASAP
Type of Contract: Permanent, full-time
About Radian Arc
Radian Arc, now part of InferX, Submer's AI cloud and GPU infrastructure platform, provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.
What impact you will have
Mission: Design, build, and operate the compute orchestration layer powering a GPU-native cloud platform for AI and high-performance workloads. (CloudStack, Kubernetes, Slurm, Argo).
The platform orchestrates GPU clusters supporting large-scale AI training and inference workloads across distributed compute infrastructure. This role bridges the current production platform, based on CloudStack, with the next-generation orchestration architecture built around Kubernetes, modern batch scheduling frameworks, and workflow orchestration systems.
You will be responsible for maintaining and evolving the existing CloudStack-based deployments while actively contributing to the design and implementation of the next-generation compute platform supporting distributed AI workloads. The role combines deep hands-on engineering with ownership of critical orchestration components, including Kubernetes-based compute orchestration, Slurm-based distributed training and batch scheduling, and workflow automation through Argo.
Working closely with networking, storage, and platform engineers, you will help implement the platform primitives that expose GPU infrastructure as a scalable, multi-tenant compute platform.
What you’ll do
CloudStack Platform Maintenance
Maintain the existing CloudStack code base used in current production deployments.
Integrate new upstream CloudStack releases into the internal platform fork.
Perform upgrades of existing customer environments to newer CloudStack versions.
Design and execute safe upgrade paths for running production environments.
Troubleshoot orchestration and provisioning issues in existing deployments.
CloudStack Networking & VPC Infrastructure
Maintain and troubleshoot CloudStack VPC networking
Work with and understand CloudStack Debian VPC routers
Manage networking implementations based on:
Open vSwitch (OVS)
OVN
Improve reliability of network orchestration components
Manage hypervisor implementations based on:
KVM
QEMU
Maintain and evolve the code responsible for QEMU GPU passthrough, including PCI mapping and exposure of L40S, RTX 6000 Pro, and H200 GPUs to virtual machines.
Next-Generation Compute Orchestration
Design orchestration and scheduling primitives for the next-generation platform based on:
Kubernetes
Slurm
Argo Workflows
Build orchestration workflows that expose GPU and CPU compute resources to platform users.
Integrate compute orchestration with storage and networking services.
Work closely with networking, storage engineers, and platform software engineers to integrate platform primitives.
Kubernetes GPU Scheduling & Cluster Orchestration
Design and implement Kubernetes-based GPU/CPU scheduling infrastructure for multi-tenant AI workloads.
Configure and maintain GPU device plugins and resource allocation mechanisms.
Implement GPU scheduling strategies including:
GPU partitioning, such as MIG
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s