Jobs and Careers
BY
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
ByteDanceSan Jose, Costa Ricafull_timeVerifiedPosted 6 Aug 2026
💰 $256,000/yr($128,000/yr – $256,000/yr)
About the role
Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation/advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It provides powerful Machine Learning computing power for internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses.<br/><br/>We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.<br/>Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.<br/>Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.<br/><br/>Job Description:<br/>- Responsible for the iteration of the underlying architecture of the large model inference engine and end-to-end GPU performance optimization, through means such as operator fusion and compilation optimization, deeply optimizing GPU memory access, computing pipeline, and Stream asynchronous scheduling, eliminating inference computing bottlenecks, improving single-card inference throughput, and reducing inference latency.<br/>- Adapt to all series of GPU/NPU hardware architectures, refine the universality of the inference engine and hardware adaptability, and build a high-performance, low-loss underlying base for large model inference.<br/>- Lead the design, development, and optimization of distributed parallel solutions for large model inference scenarios, with a focus on implementing multi-dimensional parallel strategies such as tensor parallelism (TP), pipeline parallelism (PP), sequence parallelism, and MoE expert parallelism, to address core issues such as multi-card splitting and deployment of ultra-large models, high cross-card communication overhead, load imbalance, and low parallel efficiency. <br/>- Follow up on cutting-edge technologies such as global large model inference, GPU high-performance computing, distributed parallelism, and cache optimization, benchmark against mainstream inference frameworks such as vLLM and TensorRT-LLM, complete the implementation of solutions and technological innovation, continuously iterate and optimize the performance and cost advantages of the inference system, and build the core technological barriers of the team.<br/><br/>The base salary range for this position in the selected city is $128000 - $256000 annually.<br/>
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s