Jobs and Careers
2K

Senior Site Reliability Engineer

2K
United Statesfull_timeVerifiedPosted 24 Jun 2025
💰 $145,620/yr($98,400/yr$145,620/yr)

About the role

 
#LI-Onsite 

On-Call Requirement: Yes (Periodic Rotation)

Who We Are

 2K is headquartered in Novato, California and is a wholly owned label of Take-Two Interactive Software, Inc. (NASDAQ: TTWO). Founded in 2005, 2K Games is a global video game company, publishing titles developed by some of the most influential game development studios in the world. Our studios responsible for developing 2K’s portfolio of world-class games across multiple platforms, include Visual Concepts, Firaxis, Hangar 13, CatDaddy, Cloud Chamber, 31st Union, HB Studios, and 2K SportsLab. Our portfolio of titles is expanding due to our global strategic plan, building and acquiring exciting studios whose content continues to inspire all of us! 2K publishes titles in today’s most popular gaming genres, including sports, shooters, action, role-playing, strategy, casual, and family entertainment.

 Our team of engineers, marketers, artists, writers, data scientists, producers, thinkers and doers, are the professional publishing stewards of 2K’s portfolio currently includes several AAA, sports and entertainment brands, including global powerhouse NBA®️ 2K,  renowned BioShock®️, Borderlands®️, Mafia, Sid Meier’s Civilization®️ and XCOM®️ brands; popular WWE®️ 2K and WWE®️ SuperCard franchises, TopSpin 2K25, as well as the critically and commercially acclaimed PGA TOUR®️ 2K

 At 2K, we pride ourselves on creating an inclusive work environment, which means encouraging our teams to Come as You Are and do your best work! We encourage ALL applicants to explore our global positions, even if they don’t meet every requirement for the role.  If you're interested in the job and think you have what it takes to work at 2K, we encourage you to apply!

 

What We Need

 We are seeking a Senior Site Reliability Engineer (SRE) with deep expertise in Unix/Linux systems architecture, distributed infrastructure, and automation tooling to help scale and sustain mission-critical platforms that serve millions of active users worldwide. You’ll play a leading role in building resilient, high-performance services for live gaming environments—balancing system stability, scalability, and operational velocity.


  As part of our SRE team, you’ll work across a complex technology stack spanning AWS, GCP, and hybrid on-prem environments. You’ll be responsible for building auto-scaling, self-healing Unix-based systems, optimizing OS internals, and integrating authentication across enterprise identity systems. You’ll lead the design of high-availability architecture, implement disaster recovery, apply advanced performance tuning across kernel, network, and filesystem layers, and define/enforce observability standards using Datadog, Grafana, and open-source telemetry tools. Your efforts will power real-time insights, automated alerting, and rapid incident detection and resolution. As a senior member of the on-call rotation, you’ll handle critical outages, lead post-mortems, and design long-term preventative solutions.


 Automation is foundational to this role. You’ll build and maintain infrastructure-as-code (IaC) with tools like Terraform, puppet, and Ansible, orchestrating deployments, configurations, and updates across heterogeneous environments. You’ll extend platform APIs and backend tooling using Python, and Shell scripts, driving continuous improvement in platform delivery.


 Collaboration is key: You’ll partner with backend and gameplay engineers to embed reliability into every layer of the tech stack. You’ll contribute to shared reliability standards, CI/CD integration pipelines, provisioning templates, and internal documentation. As a mentor, you’ll share your expertise in debugging, system architecture, and tooling best practices, empowering engineers across disciplines to build complex resilient systems.

 

What You’ll Do

Systems Design, Scaling & Resilience

  • Design and operate distributed Unix-based systems (Red Hat, Ubuntu, Debian, CentOS).
  • Implement auto-scaling and self-healing infrastructure to ensure uptime and durability.
  • Tune system internals including kernel parameters, networking, and filesystems for high performance.
  • Maintain timely OS patching and compliance posture across environments.
  • Integrate systems with enterprise identity services such as Active Directory, LDAP, and Kerberos.

Automation & Infrastructure as Code

  • Build and maintain infrastructure automation using Terraform, puppet, Ansible.
  • Automate deployment pipelines, service configurations, and patch management.
  • Develop scripts and services in Python, and Bash/Shell to enhance infrastructure delivery workflows.
  • Extend APIs and platform automation to d

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

2K

View company profile →