Jobs and Careers
FA

Senior Site Reliability Engineer

Fastly, Inc.
San Francisco, United Statesfull_timeVerifiedPosted 5 Feb 2024
💰 $209,740/yr($167,790/yr$209,740/yr)

About the role

Fastly helps people stay better connected with the things they love. Fastly’s edge cloud platform enables customers to create great digital experiences quickly, securely, and reliably by processing, serving, and securing our customers’ applications as close to their end-users as possible — at the edge of the Internet. The platform is designed to take advantage of the modern internet, to be programmable, and to support agile software development. Fastly’s customers include many of the world’s most prominent companies, including Vimeo, Pinterest, The New York Times, and GitHub.

We're building a more trustworthy Internet. Come join us.

Foundation Engineering at Fastly is looking for a Site Reliability Engineer to join our Cloud and Container Services team. This role is focused on helping to scale and manage Fastly’s Kubernetes based platform for control plane services. This platform is built on top of multiple public cloud services and contains many Kubernetes ecosystem components. We’re working to scale out our platform to support growth and at the same time evolve to address new business priorities. A successful candidate will help expand our platform feature set, support existing users and onboard new services, and drive efficiency while maintaining a secure platform.

What You'll Do

  • Design, build and operate infrastructure (cloud, Fastly datacenter) to enable reliable and rapid deployment, effective monitoring, and resilient operation in a large-scale Linux environment. The majority workloads are containerized but some are using native cloud services such as compute and storage
  • Diagnose and resolve performance and reliability issues across the stack: application, operating system, network, 3rd party services and APIs, including cross-application dependencies
  • Deploy and support complex 3rd party and internally developed applications
  • Write tools to automate maintenance and deployment of servers, services, and applications
  • Collaborate with internal users and continually evolve the platform and its operations using solid engineering practices
  • Drive projects sometimes independently and sometimes collaboratively across time zones
  • Configure access and manage operations within a multi-cloud environment

What We're Looking For

  • Experience running high availability systems and supporting distributed infrastructure. You have designed services with fault tolerance and across geographies.  You have deployed and managed multi-tiered services.
  • Understanding of Linux systems, high and low level. You have used tcpdump and tracing tools
  • Experience building and operating production-grade kubernetes clusters in multiple regions, clouds or data centers. You have experience with tooling in the CNCF space, such as: prometheus, flux, helm, etc.
  • Experience with programming languages such as Go and Python. You can read code and reason about what it does. You can write code within an existing large code base such as adding features. You can create medium-sized programs from scratch such as custom kubernetes controllers, custom prometheus exporters, and building tooling to help manage infrastructure
  • Experience with infrastructure and configuration management tooling. You have used Terraform to manage infrastructure
  • Experience provisioning and managing users and resources with cloud providers such as AWS and GCP. You have provisioned users and accounts using both graphical user interfaces and infrastructure as code frameworks. You have experience provisioning components on public cloud and understand how they work together in creating a multi-tiered service. You have deployed and managed services built on top of public cloud components such as EC2, S3, and GKE
  • Experience with CI/CD and GitOps tooling. You can iterate infrastructure via pipelines through code changes. You have experience using Github including creating and reviewing pull requests
  • Experience with monitoring tools such as Prometheus, Datadog and Grafana
  • Experience working on a distributed team. You have experience collaborating across timezones. You can articulate challenges of distributed teams and how you mitigate them

Work Hours: 

  • This position will require you to be available during core business hours. 

Work Locations & Travel Requirements:

This position is open to the following preferred office locations:

  • San Francisco, CA
  • Los Angeles, CA
  • Denver, CO
  • New York, NY 

Fastly currently embraces a largely hybrid model for most roles which allows employees flexibility to split their time between the office and home. 

Salary:

The estimated salary range for this position is $167,790 to $209

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Fastly, Inc.

View company profile →