Senior Platform Engineer - USDS
TikTokAbout the role
About the Team
The Cyber Defense & Engineering team is missioned to run and operate security infrastructures, platforms and technologies, as well as to support cross-functional teams to protect our users, products and infrastructures. This team is responsible for enhancing security tools and identifying vulnerabilities, with a specific focus on content assurance and the application of large language models (LLMs). You'll collaborate cross-functionally with partners inside and outside TikTok to fortify our products and users' security, helping to establish TikTok as the most trusted platform.
About the Role
We are seeking a hands-on Platform Engineer to architect, build, and operate the greenfield on-premise infrastructure that powers our next-generation AI initiatives.
This is a unique opportunity to build an AI-native platform from the ground up. You will bridge the gap between classic security services such as SPIFFE/Spire, which enables all our service-to-service communication, and modern AI workloads, leveraging a stack that includes Kubernetes, GPUs, Vector Databases, and streaming technologies like Kafka and Flink. You will solve complex challenges in distributed computing, network latency, and automated provisioning to foster a culture of innovation and velocity. There is also responsibility for designing multi-tenant cloud architecture.
Responsibilities
- Architect and operate highly available, on-premise Kubernetes-based GPU compute cluster. You will manage the scheduling and orchestration of high performance workloads. Build robust CI/CD pipelines specifically for LLM applications.
- Infrastructure as Code (IaC): Lead the design and implementation of IaC (using Terraform, Ansible or Saltstack) to fully automate the provisioning bare metal servers, network and storage layers, and ensuring the environment is reproducible and idempotent.
- Lead and perform hands-on technical work, including architecture design and code development for an on-premise, highly scalable, and parallelized infrastructure. The role includes developing internal tools to manage the entire lifecycle of a large scale RAG pipeline .
- Architect, implement, and manage a high-performance compute cluster for LLM workloads. This involves the selection and configuration of specialized hardware like GPUs, as well as the design of a robust network fabric to facilitate efficient inter-node communication for parallel processing.
- Implement security best practices for a private data center environment. This includes configuring network firewalls, managing access controls, and encrypting data at rest and in transit.
- Establish comprehensive monitoring and alerting systems to track the health and performance of the compute cluster and LLM workloads. This involves analyzing metrics related to GPU utilization, memory usage, network throughput, and model inference latency. You will proactively resolve performance issues to enhance platform reliability and operational support for internal teams.
- Collaborate with internal stakeholders to optimize resource utilization and improve the platform's efficiency. You'll work closely with data scientists and machine learning engineers to understand their compute needs and ensure the infrastructure is optimized for their specific workloads.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s