Jobs and Careers
BY
Research Scientist - ByteBrain - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
ByteDanceSan Jose, Costa Ricafull_timeVerifiedPosted 4 Aug 2026
💰 $387,600/yr($218,400/yr – $387,600/yr)
About the role
ByteBrain is ByteDance’s AI for Infrastructure (AI4Infra) platform, dedicated to improving the efficiency, reliability, and intelligence of large-scale infrastructure systems through AI and machine learning. ByteBrain supports a wide range of infrastructure domains, including AI data center supply chains, AIOps, Operations Research and AgentOps, powering infrastructure optimization at massive scale.<br/><br/>Why This Role Is Unique<br/>This role sits at the intersection of: Operations Research × AIOps × AI for Infra<br/>You will have the opportunity to solve some of the most challenging optimization problems behind large-scale AI datacenters while pioneering the next generation of AI-powered decision-making systems, where LLMs, and optimization algorithms work together to improve efficiency, resource utilization, and operational intelligence across ByteDance's global infrastructure.<br/><br/>Responsibilities<br/>- Design and develop AI, machine learning, and optimization algorithms to improve the efficiency, reliability, and performance of large-scale infrastructure systems and AI supply chain. Areas may include AIOps, operations research, software engineering, AgentOps, and system optimization.<br/>- Drive the deployment, scaling, and continuous improvement of algorithms in production environments, supporting large-scale services.<br/>- Identify optimization opportunities and emerging challenges from real-world infrastructure scenarios, translating them into impactful research and engineering solutions.<br/>- Conduct cutting-edge research and publish high-quality papers in top-tier conferences and journals.<br/><br/>Topic Content: <br/> With the large-scale adoption of LLMs and AI agents, traditional cloud-native infrastructure can no longer meet the ultra-high performance and elasticity requirements of AI workloads. This topic conducts systematic research across the entire AI infrastructure stack:<br/>1. Network and Observability: Research intelligent fault localization and root cause analysis for large-scale AI clusters, combined with intelligent tuning of time-series databases to improve cluster stability.<br/>2. Storage Systems: Develop serverless high-performance elastic file systems and storage acceleration architectures specifically for AI scenarios, explore hardware-software co-optimization for DPU, and overcome AI storage performance bottlenecks.<br/>3. Data Center Power Scheduling: Research GPU/CPU/MEM heterogeneous collaborative scheduling technologies, build a heterogeneous power orchestration system for AI agents, and address scheduling challenges including heterogenous workloads and state dependencies.<br/>4. Vector Retrieval: Optimize core vector retrieval technologies for LLM-powered applications, building a cloud-native distributed vector index engine to meet ultra-large-scale vector retrieval demands with low latency and low cost.<br/>5. Intelligence and Agent Architecture: Explore automatic infrastructure optimization based on AI Agent workflows, build a self-evolvable business agent framework, and enable full-stack intelligent optimization through AI for Infra. <br/><br/>This topic aims to build a next-generation AI-native infrastructure to support the deployment of LLMs and AI agents, improve resource utilization, reduce costs, support elastic scaling, and drive the technological evolution of AI infrastructure.<br/><br/>The base salary range for this position in the selected city is $218400 - $387600 annually.<br/>
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s