Jobs and Careers
RE

Staff SRE/DevOps Engineer (Platform Reliability & Security)

Red Cell Partners
United Statesfull_timeVerifiedPosted 25 Mar 2026
💰 $230,000/yr($170,000/yr$230,000/yr)

About the role

About Us

Red Cell Partners is an incubation firm building and investing in rapidly scalable technology-led companies that are bringing revolutionary advancements to market in three distinct practice areas: healthcare, cyber, and national security. United by a shared sense of duty and deep belief in the power of innovation, Red Cell is developing powerful tools and solutions to address our Nation’s most pressing problems.

About Trase

Co-founded in 2023 by Joe Laws and Grant Verstandig, Trase Systems is AI, Uncomplicated. Trase empowers enterprise leaders to harness the full potential of AI without the associated complexity and risks. We are an end-to-end solution for deploying, managing, and optimizing AI in the enterprise. Our platform specializes in bridging the “last mile” of AI adoption, unlocking AI's full potential while driving efficiency and significant cost savings. Trase is at the forefront of AI Agent innovation, topping the Hugging Face GAIA Leaderboard for Generalized AI Assistants, ahead of industry giants such as Google, Meta, Microsoft, and OpenAI. We are leveraging our cutting-edge technologies to develop mission-critical agentic applications in complex industries such as Healthcare, Oil & Gas, and National Security. 

About the Role

Location: Seattle, WA area

As a Staff DevOps Engineer, you will own the reliability, security, and operational foundations of Trase OS, the shared platform that powers every Trase deployment.

This is a core engineering role, not a support function. You will design and operate the infrastructure, delivery systems, and runtime controls that allow the OS platform  to safely run long-lived, multi-step workflows under real security and compliance constraints.

Your work directly shapes the architecture of the platform and determines how confidently Trase can scale.

Why this role is needed

Trase OS is a distributed system with long-lived, stateful workflows and strict security constraints. Reliability and security are core architectural concerns, not operational afterthoughts.

Without strong infrastructure ownership, small failures can become systemic instability, and scaling introduces risk instead of leverage.

This role exists to:

  • Prevent systemic instability at the platform level
  • Establish reliability and security as first-class design properties
  • Enable safe, repeatable scaling as customer count, workload complexity, and regulatory expectations grow

What you'll do

You will work on:

  • A shared platform that powers all Trase deployments
  • Long-running, multi-step workflows that must survive failures and restarts
  • Security-first architecture including authentication, RBAC, auditability, and traceability
  • Platform-level observability so customers can trust what the system did and why
  • Infrastructure and delivery systems that turn one-off builds into reusable platform capabilities

You will:

  • Own deployment, runtime reliability, and security for Trase OS services and infrastructure
  • Design and operate cloud infrastructure supporting secure, repeatable multi-environment deployments
  • Build and maintain CI/CD systems, release orchestration, and environment management to ensure safe, predictable delivery
  • Own observability systems (metrics, logs, traces, alerting) enabling rapid detection, diagnosis, and recovery
  • Design and operate networking and traffic management, including secure service-to-service communication and rollout patterns
  • Implement and operate policy enforcement mechanisms (e.g., service mesh controls, authentication/authorization integration, runtime guardrails)
  • Define, instrument, and operate service level objective and indicators (SLOs/SLIs) and error budgets, and use them to drive engineering decisions
  • Ensure the system is resilient by design, including:
    • Failure isolation and blast-radius control
    • Safe retries and idempotency
    • State recovery for long-lived workflows
    • Capacity planning and operational runbooks

Staff-level technical leadership

  • Lead infrastructure and reliability architecture across teams building on Trase OS
  • Set standards for production readiness, security posture, and operational excellence
  • Drive adoption of SLO-driven engineering practices across the platform
  • Partner with platform, product, and DevEx engineers to align architecture with developer velocity and customer trust
  • Mentor engineers (including se

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Red Cell Partners

View company profile →