Site Reliability Engineer
Dutch Bros CoffeeAbout the role
It's fun to work in a company where people truly believe in what they are doing. At Dutch Bros Coffee, we are more than just a coffee company. We are a fun-loving, mind-blowing company that makes a difference one cup at a time.
Position Overview:
As a Site Reliability Engineer (SRE), you will combine software engineering principles with systems administration to help build and run scalable, reliable, and secure systems. You’ll support the implementation of tools and processes that improve service reliability, reduce manual toil, and enhance observability across our multi-cloud enterprise. This role is highly collaborative, working closely with Platform Development, IT Operations, DevOps, and Product teams to deliver performant, highly available systems and services. You’ll participate in incident response and root cause analysis, continuously improve infrastructure and deployment practices, and help support our SLOs and SLIs across critical systems.
Job Qualifications:
Bachelor’s degree in Computer Science, Engineering, or related field—or equivalent work experience
3+ years of SRE, DevOps, or systems engineering experience
Proficient in at least one programming language (e.g., Python, Java, Golang)
Strong understanding of platform systems, networking, and systems administration tools
Hands-on experience with public cloud providers (AWS and/or Azure)
Familiarity with CI/CD tools and pipelines (e.g., GitHub Actions, Jenkins)
Experience with Infrastructure as Code tools like Terraform or Ansible
Working knowledge of observability stacks and performance monitoring
Experience with containerization and orchestration (Docker, Kubernetes preferred)
Experience supporting compliance-driven infrastructure requirements including SOC 2, PCI DSS, and public company regulatory needs (e.g., SOX)
Understanding of incident response best practices and on-call support
Excellent problem-solving, documentation, and communication skills
Ability to collaborate across engineering, security, and compliance teams
Location Requirement:
This role is located in Tempe, Arizona. This position is required to be in office 4 days per week (Mon-Thurs); Fridays are optional remote work days.
Key Result Areas (KRAs):
Support the development and improvement of systems to meet reliability, availability, and security objectives
Help design and build automated solutions that reduce operational overhead while supporting compliance requirements such as SOC 2, PCI DSS, and SOX
Implement and maintain monitoring, alerting, and observability solutions to meet both operational and compliance standards
Participate in incident response and assist with troubleshooting, diagnostics, and resolution
Contribute to root cause analysis, post-incident reviews, and ensure implementation of preventative measures aligned to regulatory requirements
Collaborate with security, compliance, and infrastructure teams to maintain alignment with public company data governance and audit requirements
Support scaling efforts by anticipating capacity, performance, and regulatory considerations
Create and maintain runbooks, technical documentation, and SRE processes that support audit-readiness
Continuously identify and resolve recurring reliability and compliance-related issues
Other duties as assigned
Must be able to collaborate in-person with occasional impromptu in-person meetings
Skills:
Observability and Monitoring Tools (e.g., Datadog, Prometheus, Grafana)
Programming/Scripting (e.g., Python, Bash)
CI/CD Pipelines and Automation
Systems Administration (Linux, Networking)
Infrastructure as Code (Terraform, Ansible)
Containerization (Docker, Kubernetes)
Problem-solving, Incident Management, and Root Cause Analysis
Clear, Concise Technical Documentation
Understanding of compliance frameworks including SOC 2, P
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s