Back to ByteDance jobs
B

Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

Singapore
RegularBackend

Job Description

Team Introduction Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation/advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It provides powerful Machine Learning computing power for internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses.

We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.

Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.

Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to ByteDance and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early. Responsibilities

  • Responsible for the overall architecture design and implementation of model inference services, building a high-performance, highly available, and scalable enterprise-level inference system for large-parameter, high-complexity AI models, overcoming various architectural challenges in the implementation of complex model inference, and supporting the efficient launch of models across all business scenarios.
  • Responsible for the R&D and optimization of the core modules of the inference framework, covering core capabilities such as inference engine scheduling, monitoring and alerting, canary release, etc., continuously iterating on the framework performance, and resolving performance bottlenecks, resource bottlenecks, and stability issues in high-concurrency and large-model inference scenarios.
  • Keep track of the latest inference technologies in the industry, conduct technology selection and innovation in combination with business scenarios, accumulate distributed high-concurrency service architecture solutions, and promote the upgrade and standardization of the team's technical system.

Qualifications Minimum Qualifications:

  • Individuals who are completing or have recently completed a Bachelor's/ Master's degree in computing or a related discipline.
  • Familiar with basic Linux commands, with solid C/C++ programming skills and knowledge of data structures and algorithms
  • Familiar with the basic principles of multi-threaded concurrency, proficient in basic usages such as thread usage, synchronization locks, and thread pools, able to identify common concurrency issues, and possess the ability to perform basic performance tuning in multi-threaded scenarios
  • Have experience in R&D projects of high-concurrency distributed services, and be familiar with service latency and resource optimization
  • Possess good learning and execution abilities, be willing to proactively understand the model inference service architecture, and have problem analysis and abstraction capabilities
  • Possess good cross-team collaboration skills, communication and presentation skills, and document writing skills, have strong sense of responsibility and stress tolerance, and be able to drive the resolution of complex technical issues and the implementation of projects

Preferred Qualifications:

  • Have practical project experience and understanding of source code in high-concurrency services/frameworks such as Redis, RocksDB, BRPC, GRPC, etc.
  • Understand the operating mechanism of GPUs, and have relevant project experience and optimization capabilities in GPU service resource management and control

About ByteDance

First seen: August 3, 2026
Last updated: August 5, 2026