Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Job Description
Team Introduction Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation/advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It provides powerful Machine Learning computing power for internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses.
We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.
Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to ByteDance and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities -Responsible for the iteration of the underlying architecture of the large model inference engine and end-to-end GPU performance optimization, through means such as operator fusion and compilation optimization, deeply optimizing GPU memory access, computing pipeline, and Stream asynchronous scheduling, eliminating inference computing bottlenecks, improving single-card inference throughput, and reducing inference latency. -Adapt to all series of GPU/NPU hardware architectures, refine the universality of the inference engine and hardware adaptability, and build a high-performance, low-loss underlying base for large model inference. -Lead the design, development, and optimization of distributed parallel solutions for large model inference scenarios, with a focus on implementing multi-dimensional parallel strategies such as tensor parallelism (TP), pipeline parallelism (PP), sequence parallelism, and MoE expert parallelism, to address core issues such as multi-card splitting and deployment of ultra-large models, high cross-card communication overhead, load imbalance, and low parallel efficiency. -Follow up on cutting-edge technologies such as global large model inference, GPU high-performance computing, distributed parallelism, and cache optimization, benchmark against mainstream inference frameworks such as vLLM and TensorRT-LLM, complete the implementation of solutions and technological innovation, continuously iterate and optimize the performance and cost advantages of the inference system, and build the core technological barriers of the team.
Qualifications Minimum Qualifications:
- Individuals who are completing or have recently completed a Bachelor's/ Master's degree in computing or a related discipline.
- Solid foundation in computer low-level knowledge, proficient in C/C++ and Python programming, skilled in CUDA programming and familiar with GPU hardware architecture principles, and well-versed in GPU memory models, computing scheduling, and communication mechanisms;
- Proficiently master the underlying development and implementation of various basic operators in Deep learning, be well-versed in GPU adaptation and optimization of core operators such as matrix operations, normalization, and activation functions, and be able to independently complete operator handwritten reconstruction, memory access optimization, vectorization acceleration, and precision alignment to ensure high performance and high stability of operator inference.
- Familiar with the end-to-end process of deep learning inference compilation, understand core compilation technologies such as computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation, and be able to simplify the inference process, reduce GPU memory usage, and decrease inference latency through compilation-level improvements, thereby significantly enhancing the throughput efficiency of model inference.
- Proficient in using GPU performance analysis tools such as Nsight and Profiler, able to accurately identify performance bottlenecks such as computing power waste, memory access blockage, and scheduling redundancy during the inference process, possess the thinking of software-hardware collaborative optimization, capable of outputting systematic optimization solutions and completing implementation iterations, and adaptable to the requirements of industrial-level high-concurrency, low-latency inference business.
- Possess good cross-team collaboration skills, communication and presentation skills, and document writing skills, have strong sense of responsibility and stress tolerance, and be able to drive the resolution of complex technical issues and the implementation of projects.
Preferred Qualifications:
- Thoroughly understand the core principles of large model inference, proficiently master the core technologies of model parallelism, have experience in implementing distributed inference solutions such as tensor parallelism, pipeline parallelism, and sequence parallelism, and be familiar with multi-card communication, load balance, and parallel efficiency optimization methods.
- Those with experience in secondary development and Performance optimization of mainstream large model inference frameworks such as vLLM, SGLang, TensorRT-LLM, etc. are preferred.