Jobs and Careers
GI

Intermediate Site Reliability Engineer, Tenant Scale: Tenant Services

GitLab
Remote, Americas; Remote, EMEARemotefull_timeVerifiedPosted 9 Jan 2026

About the role

GitLab is an open-core software company that develops the most comprehensive AI-powered DevSecOps Platform, used by more than 100,000 organizations. Our mission is to enable everyone to contribute to and co-create the software that powers our world. When everyone can contribute, consumers become contributors, significantly accelerating human progress. Our platform unites teams and organizations, breaking down barriers and redefining what's possible in software development. Thanks to products like Duo Enterprise and Duo Agent Platform, customers get AI benefits at every stage of the SDLC. 

The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.

An overview of this role

As a Site Reliability Engineer (SRE) at GitLab, you keep GitLab.com and other production systems running smoothly for millions of users by combining pragmatic operations with strong software engineering practices. You focus on the systems layer (operating systems, storage, networking) and edge services and Kubernetes workloads, designing and operating highly scalable, reliable, and secure infrastructure that supports one of the largest single-tenancy open source SaaS sites on the Internet. You’ll work across the Infrastructure organization to automate away toil, improve availability and performance, and respond to incidents during your local daytime hours as part of a globally distributed on-call rotation. In this role, you’ll help Tenant Services safeguard and scale customer data while increasing automation so GitLab can continue to grow with enterprise-level expectations for reliability and availability.

What you’ll do

  • Design and implement highly scalable infrastructure for GitLab.com to support current and future growth.
  • Collaborate with cross-functional teams across the Infrastructure organization to plan and deliver projects that shape GitLab’s platform direction.
  • Operate and improve edge services and Kubernetes workloads, acting as a subject matter expert within the infrastructure department.
  • Participate in a global on-call rotation during your local daytime hours, respond to production incidents, and contribute to clear, constructive incident reviews.
  • Reduce toil by automating operational tasks and building tools that improve reliability, availability, and scalability.
  • Apply infrastructure as code and configuration management practices to manage cloud resources and environments consistently.
  • Write and maintain production-quality code, preferably in Go or Ruby, to enhance our systems and automation toolchain.

What you’ll bring

  • Background working with the Kubernetes ecosystem, including tools such as Helm, and running production workloads.
  • Experience operating cloud infrastructure on platforms like Google Cloud Platform or Amazon Web Services, especially networking, hosted Kubernetes services, and scaling.
  • Hands-on practice with infrastructure as code and configuration management tools such as Ansible or Chef.
  • Strong programming skills in a modern language, preferably Go or Ruby, applied to automation and reliability problems.
  • Ability to clearly define problems, think beyond short-term fixes, and design solutions that improve systems over time.
  • Consistent focus on reducing toil through automation and thoughtful system design.
  • Independent, proactive working style with a bias for action and comfort operating as a “manager of one” in a distributed, asynchronous environment.
  • Clear written and verbal communication skills, with openness to candidates who bring transferable experience from related reliability, infrastructure, or platform roles.

About the team

Tenant Services is the team responsible for safeguardi

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

GitLab

View company profile →