
AI Infrastructure Engineer Intern (Algorithm Infrastructure) - 2027 Start (PhD)
Job Description
TikTok is an international short video platform available in over 150 countries and regions, where we aim to inspire creativity and bring joy by helping people discover authentic and interesting moments. TikTok has offices around the world, with global headquarters in Los Angeles and Singapore, and additional offices in New York, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo, among others.
We are dedicated to building the inference infrastructure for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems. Our mission is to provide a robust, scalable, and high-performance foundation for distributed serving, heterogeneous scheduling, and low-latency inference at massive scale. You will work on some of the most challenging problems in large-model online serving, spanning traffic orchestration, throughput and latency optimization, kernel efficiency, and production reliability for next-generation AI systems.
We are looking for talented individuals to join us for an internship. PhD internships at Our Company provide students with the opportunity to actively contribute to our products and research, as well as to the organization's future plans and emerging technologies. Our dynamic internship experience blends hands-on learning, enriching community-building and professional development events, and collaboration with industry experts. Applications will be reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume (Start date, End date).
Responsibilities:
- Build and evolve next-generation inference systems for large-scale online traffic, including global scheduling across heterogeneous compute resources, high-concurrency load balancing, and efficient batch formation
- Optimize distributed inference for 200B+ models and complex multimodal models through TP, EP, DP, and related strategies to improve throughput and latency in production
- Develop high-performance kernels for frontier model architectures such as MoE, emerging attention mechanisms, and multimodal fusion layers using CUDA, Triton, and related tools
- Explore AI-driven infrastructure for inference systems, including AI Agents for kernel optimization, performance tuning, consistency validation, deployment pipelines, and intelligent operations
Qualifications Minimum Qualifications:
- Currently pursuing a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline in Computer Science, Software Engineering, Artificial Intelligence, Mathematics, or related fields
- Familiarity with large-model architectures and strong system design skills for complex, high-concurrency environments
- Strong engineering skills in performance optimization and production system development
Preferred Qualifications:
- Strong understanding of asynchronous scheduling, resource pooling, and load balancing in distributed microservice systems
By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://careers.tiktok.com/legal/privacy