Site Reliability Engineer (SRE)
AbridgeAbout the role
Abridge was founded in 2018 with the mission of powering deeper understanding in healthcare. Our AI-powered platform was purpose-built for medical conversations, improving clinical documentation efficiencies while enabling clinicians to focus on what matters most—their patients.
Our enterprise-grade technology transforms patient-clinician conversations into structured clinical notes in real-time, with deep EMR integrations. Powered by Linked Evidence and our purpose-built, auditable AI, we are the only company that maps AI-generated summaries to ground truth, helping providers quickly trust and verify the output. As pioneers in generative AI for healthcare, we are setting the industry standards for the responsible deployment of AI across health systems.
We are a growing team of practicing MDs, AI scientists, PhDs, creatives, technologists, and engineers working together to empower people and make care make more sense. We have offices located in the SoHo neighborhood of New York, the Mission District in San Francisco, and Lawrenceville in Pittsburgh.
The Role
Abridge’s services and engineering team are in hyperscale mode. We are looking for experienced SREs to join our team and help scale all of our infrastructure, developer experience, and site reliability in kind. You’ll work on a centralized Platform team whose work spans platform building, adoption, and ongoing support of existing tooling and software. You will partner with other teams and may embed with them for weeks or months. The platform we are building needs to maximize both engineering velocity and security, will be under tremendous scale, and presents many opportunities to leverage creativity, autonomy, and leadership to take things 0 to 1.
This is a unique opportunity in the industry to rapidly grow your career in a rapidly growing company leveraging the best of emerging technologies.
What You'll Do
Design and implement build pipelines, branching strategies, and release management tooling that will serve an engineering team that is doubling in size and massively growing the volume of code that is being shipped and that must be tested.
Build out our observability platform as the systems continue to grow, enabling logs, metrics, and distributed tracing at scale, as well as building out dashboards with golden metrics and error budgets that allow us to keep a pulse on the system’s growth.
Partner with teams to leverage observability tooling and observability driven development to identify bottlenecks in performance and availability and fix them.
Design and implement cloud security tooling and processes that strike an effective balance between engineering velocity and software and data security.
Help advocate for, design, implement, and adopt fast and scalable application testing pipelines including end to end UI tests as well as hyperscale load tests.
Help build an entirely new cloud stack that is fully provisioned through IaC and highly secure; leverage this to build ephemeral environments and multi-tenant/multi-region production deployments.
Uplevel our ability to respond to incidents by improving observability, runbooks, and incident response muscle across the organization.
Bridge the gap between local development and production environments in a way that is seamless for engineers and maximizes engineering velocity and security while minimizing quality issues arising from environment drift and configuration tangles.
Evangelize, document, and train the engineering team on the solutions being built and uplevel them on cloud native design strategies and tools.
Be a public evangelist for Abridge in the global platform engineering community, including conferences, open source, and research as we pioneer new AI-first cloud-native-first security-first implementations at scale.
Who You Are
6+ years of software engineering experience focused on distributed systems or tooling, with an interest in engineering enablement and software scaling.
At least 2 years experience as a back-end engineer focused on system performance and scalability.
Experience building on Kubernetes and scaling compute services on Kubernetes; experience with related cloud native technologies including ArgoCD, Argo Rollouts, Istio, etc.
Experience (or strong interest in) creating and maintaining CI/CD pipelines for both Infrastructure as code deployments as well as application code deployments. (Terragrunt, Atlas, ArgoCD, Octopus Deploy, Travis CI, etc.)
Comfortable implementing and securing services in Google Cloud Platform with Infrastructure as Code, including GCP Projects, VPC Networks, Google Kubernetes Engine, and IAM Roles, Groups and policies. Cand
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s