Staff SRE - Data Reliability
Fastly, Inc.About the role
Fastly helps people stay better connected with the things they love. Fastly’s edge cloud platform enables customers to create great digital experiences quickly, securely, and reliably by processing, serving, and securing our customers’ applications as close to their end-users as possible — at the edge of the Internet. The platform is designed to take advantage of the modern internet, to be programmable, and to support agile software development. Fastly’s customers include many of the world’s most prominent companies, including Vimeo, Pinterest, The New York Times, and GitHub.
We're building a more trustworthy Internet. Come join us.
Posting Open Date: Sept. 8, 2025
Anticipated Posting Close Date*: Oct. 6, 2025
*Job posting may close early due to the volume of applicants.
Staff Site Reliability Engineer - Data Reliability
The Data Reliability team is looking for a talented Staff Site Reliability Engineer to help build and support the next generation of data stores for Fastly. The ideal candidate will have experience working with backend and data services in both cloud and physical systems; be skilled with configuration management tools such as Terraform; and be able to develop internal administration tools in Go and similar. Our team is responsible for supporting the infrastructure, orchestration, and reliability needs of some of Fastly’s most data-intensive applications, using technologies like Terraform, Elasticsearch, ClickHouse, Prometheus, MySQL, and Redis in both cloud- and hardware-based environments. Our systems directly contribute to our customers' success by providing our product teams with a platform for effective and reliable delivery of high-quality, high-throughput, globally distributed data systems and products. You will be integral to this mission. We are a distributed team, and employ a variety of styles to get our work done—though we put a high value on working collaboratively, we also rely on asynchronous communication.
What You'll Do:
- Lead full lifecycle projects from design and development through roll out and maintenance
- Deploy and maintain several different types of critical data storage systems on scales from gigabytes to petabytes
- Develop statistics and dashboards to measure service-level objectives for these systems
- Maintain and create tools for management of configuration, backup, and authenticated access to data systems using peer review, CI/CD, and both daemon- and container-based deployment
- Write code that is performant, maintainable, clear, and concise and contribute to code reviews, improving the codebase and other team processes
- Technical leadership of full lifecycle projects, driving project progress and collaborating with project stakeholders - Coordinate and communicate with the team members and across other technical and cross functional teams. Foster relationships with other teams to understand and provide for end-user needs.
- Help project, plan, and scale for growth
- Document processes and requirements for both intra- and cross-team knowledge transfer
- Mentor and support other engineers, fostering a culture of knowledge sharing, innovation, and collaboration within the team
- Participate in on-call rotation as needed
What We're Looking For:
- Significant professional experience building and maintaining at least one category of backend system, including both APIs and data stores (relational-, column-, vector-, or document-oriented) Most Staff Engineers at Fastly have more than 7 years of related experience.
- Experience measuring customer-focused performance using tools like Prometheus or DataDog
- Linux system skills, including file systems, networking, I/O analysis, and general kernel tuning
- A strong grasp of networking, routing, network protocols, and related concepts
- Advanced Terraform skills, including working with modules and remote state
- Experience developing automation tools in Go and other languages, with an aptitude for learning new languages and technologies
- Experience with container technology, Kubernetes/Docker/Helm, or similar
- A collaborative mindset with experience working across cross-functional teams, fostering a culture of respect, collaboration, and knowledge sharing
- A great teammate: communicative, collaborative, empathetic with a thoughtful, customer-driven approach
We’ll be super impressed if you have experience in any of these:
- Performance measurement using eBPF, flame graphs, etc.
- Creation of Kubernetes operators
- Leveraging a service mesh like Istio or Linkerd for service discovery and ACLs
- Knowledge of one or more host management t
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s