Principal Platform Engineer — Kubernetes & Cloud Infrastructure
OmbudAbout the role
<ul> <li><strong>Location: </strong>Denver, CO (hybrid — Tue/Wed/Thu in office)</li> <li><strong>Reports to: </strong>CEO</li> </ul> <h2><strong>The role</strong></h2> <p>Ombud's platform runs production AI workloads for enterprise customers, and we're scaling toward a self-service motion where customers onboard, ingest content, and operate the product without manual implementation. That requires an infrastructure foundation that can handle multi-tenant scale, high reliability, and the unique demands of generative AI workloads — without ballooning the AWS bill.</p> <p>We're hiring a Principal Platform Engineer to own that foundation. This is a senior individual contributor role with broad architectural authority. You will not have direct reports. You will set the technical direction for our cloud infrastructure, partner with engineering on production scaling decisions, and operate the platform with the discipline a SOC 2 / ISO 27001 customer base requires.</p> <h2><strong>What you'll own</strong></h2> <ul> <li>Production Kubernetes (EKS) clusters: capacity planning, node group strategy, gen-AI workload isolation, blast-radius containment.</li> <li>AWS infrastructure end-to-end: RDS, DMS, Kafka (MSK), ECR, networking, IAM, multi-region deployments (including Ireland for EU data residency).</li> <li>Infrastructure-as-code in Terraform — modules, environments, drift management, peer review.</li> <li>CI/CD pipelines (Jenkins, GitHub Actions, or your recommended replacement) — fast, reliable, secure builds for backend and frontend services.</li> <li>Observability: Grafana dashboards, Prometheus metrics, log pipelines, on-call alerting, SLO definition.</li> <li>Cost optimization. AWS spend is one of our top three variable costs. Reducing it by 20% is a tangible objective for this seat.</li> <li>Security posture: secrets management (Consul/Vault), IAM hygiene, vulnerability patching, support for SOC 2 and ISO 27001 audit cycles.</li> <li>Architecture leadership on the self-service infrastructure roadmap: how we onboard a customer without human intervention and scale to 10x our current tenant count.</li> <li>Documentation and runbooks that let the rest of the engineering team operate the platform when you're unavailable.</li> </ul> <h2><strong>Must-haves</strong></h2> <ul> <li>8+ years of platform, infrastructure, SRE, or DevOps experience, with at least 3+ years operating production Kubernetes at scale.</li> <li>Deep AWS expertise across compute, storage, networking, data services, and IAM.</li> <li>Production fluency with Terraform, Docker, Linux, and CI/CD systems.</li> <li>Track record of architectural decisions
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s