Senior Site Reliability Engineer, Network and Traffic - USDS
TikTokAbout the role
The infrastructure team of US Tech Services Department at TikTok supports the company's fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable.
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed infrastructures. Our SREs are tasked to ensure the infrastructure services are reliable, fault-tolerant, efficiently scalable and cost-effective.
Responsibilities
- Design, build, operate and optimize TikTok network, not limited to, backbone, public cloud and Edge/CDN.
- Design, manage, operate and contribute to the development of our Layer7 Load Balancer, CDN origin and DNS.
- Troubleshoot network and / or traffic issues in partnership with our Data Center Partners.
- Work with cross-functional teams including but not limited to compute, storage, database, application teams to drive the innovation and evolution of TikTok network.
- Work with external vendors and ISPs for device and carrier selection, perform testing and verification.
- Participate in the global oncall rotation to support production traffic and networking issues.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s