Jobs and Careers
OP

Staff Software Engineer, Site Reliability (SRE)

Optimal Dynamics
UKRemotefull_timeVerifiedPosted 28 Aug 2025
💰 $200,000/yr($160,000/yr$200,000/yr)

About the role

About Our Company

Built on over four decades of pioneering research at Princeton University, our platform represents the leading edge of innovation in freight and transportation planning. We help customers unlock double-digit revenue gains and drive smarter, data-driven operations at scale.
With the recent close of our Series C funding round led by Koch Disruptive Technologies, we’re entering an exciting new phase of growth. Today, Optimal Dynamics is a high-growth company of ~70 employees, backed by top-tier investors including Bessemer Venture Partners, The Westly Group, Activate Capital, and Koch. 

We're on a mission to redefine the way logistics decisions are made—and we’re just getting started.

About Our Team

We are a team of bright, kind, and solution-oriented people focused on creating value for our customers. We can solve problems individually, but understand that the best solutions are found when the team brainstorms ideas together. We are excited about balancing the need to deploy new solutions quickly and designing solutions that are secured, reliable, maintainable, and scalable for the long run.

About the Role 

We’re hiring a Staff Software Engineer, Site Reliability to lead reliability across our production platform. As a Staff‑level Individual contributor, you will drive strategy and hands‑on execution across incident response, SLO/SLI programs, and production readiness, directly owning highly available services in AWS; all while partnering with Platform/Infra to build paved‑road tooling in our monorepo.

This is a full‑time, remote‑friendly role open to candidates across the United States. For those who prefer an in‑office experience, our HQ in New York City offers a collaborative environment.

What You’ll Do

Reliability (≈50%)

  • Own the company‑wide incident lifecycle: standards for detection, escalation, incident command, customer comms, and high‑quality postmortems with action tracking.
  • Define and drive SLIs/SLOs for core services; build guardrails and dashboards that make reliability visible and actionable.
  • Lead production readiness reviews, capacity/performance planning, load testing, disaster recovery exercises, and resilience engineering (failure testing/chaos where appropriate).
  • Level‑up on‑call: right‑sizing rotations, paging hygiene, runbooks, auto‑remediation, and continuous improvement of MTTA/MTTR.

Security (≈30%)

  • Embed security into the delivery pipeline: dependency and image scanning, least‑privilege/IAM baselines, secrets management, and service‑to‑service auth.
  • Partner with Engineering leadership to maintain SOC 2‑aligned controls as code; make audit‑friendly evidence generation part of everyday engineering.
  • Drive secure‑by‑default patterns in the platform (e.g., network posture, data protection, runtime policies) without slowing down developers.

Platform & DevEx (≈20%)

  • Build and evolve paved roads for deploys, config, and runtime operations in our monorepo (Bazel) and CI/CD (AWS CodePipeline/CodeBuild).
  • Partner with product teams to make the “secure, reliable default” the easiest path—templates, tooling, libraries, and automation.
  • Improve observability end‑to‑end (traces, logs, metrics, alerts).

Who You Are

  • Experienced: Staff‑level IC who has led reliability programs at meaningful scale and owned incident response standards.
  • Technically Grounded: Deep, hands-on experience with infrastructure at scale, cloud, containerization, and more::
    • AWS (multi‑service)
    • ECS and/or Kubernetes containerization workloads 
    • CICD & IaC (Terraform) 
    • Production Networking/Fundamentals
  • Python Proficient: You can read/review service code and land operational improvements.
  • Data Driven: In your approach to SLOs, capacity, performance, and cost efficiency with strong observability chops
  • Influential: Able to shape direction and create simple, durable standards
  • Communicative: Excels in both technical and interpersonal communication, with strong written and verbal skills

Nice To Have (Bonus Points)

  • Aware of FinOps (cost attribution, efficient scaling) and DR/BCP program experience.
  • Familiar with secure SDLC, threat modeling, and compliance automation in a SOC 2 context.
  • Experience collaborating with Data Science/ML team

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Optimal Dynamics

View company profile →