Jobs and Careers
TI

Site Reliability Engineer - USDS (Multiple Positions)

TikTok
Mountain View, United Statesfull_timeVerifiedPosted 20 Feb 2025

About the role

About TikTok U.S. Data Security

TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. U.S. Data Security (“USDS”) is a subsidiary of TikTok in the U.S. This new, security-first division was created to bring heightened focus and governance to our data protection policies and content assurance protocols to keep U.S. users safe. Our focus is on providing oversight and protection of the TikTok platform and U.S. user data, so millions of Americans can continue turning to TikTok to learn something new, earn a living, express themselves creatively, or be entertained. The teams within USDS that deliver on this

commitment daily span across Trust & Safety, Security & Privacy, Engineering, User & Product Ops, Corporate Functions and more.



Why Join Us

Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.

Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.

To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.

At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.

Join us.



About the Team

Our team plays a crucial role in ensuring the company’s success. We seek people who are willing to learn and put in the effort to solve problems. Our challenges are not your regular day-to-day problems - you’ll be part of a team that’s developing new solutions to new challenges. It’s working fast at scale, and we’re making a difference. We are looking for talents to join us on this exciting journey!



Responsibilities

Provide site reliability engineering support to ensure highest level of availability of large-scale, fault-tolerant systems.

Deliver tools/software to improve the reliability, scalability and operability of services, including designing, developing and deploying automation to sustainably scale with quality.

Measure and monitor availability, latency and overall service health.

Practice sustainable incident response and postmortems, performing root cause analysis of incidents to influence future product design and response activities.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

TikTok

View company profile →