Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

BYTEDANCE PTE. LTD.

Singapore

On-site

SGD 100,000 - 160,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance's Data AML team builds training and inference systems for recommendation and advertising. This role is for graduates who want to tackle challenging GPU-accelerated workloads and join a world-class ML platform supporting Douyin, Jinri Toutiao, and Xigua Video.

You will work on optimizing the large-model inference engine, exploring tensor and pipeline parallelism, and benchmarking against leading frameworks.

Qualifications

  • Bachelor's or Master's degree in computing or related field.
  • Solid foundation in low-level computing, proficient in C/C++ and Python, CUDA programming, GPU architecture knowledge.
  • Proficient in deep learning operators and GPU optimization, including memory access and scheduling.
  • Familiar with deep learning inference compilation: graph optimization, fusion, quantization.
  • Experience with GPU performance tools (Nsight, Profiler) and software-hardware co-optimization.
  • Strong collaboration, communication, and project execution skills.

Responsibilities

  • Iterate the architecture of the large model inference engine to optimize GPU performance and reduce latency.
  • Adapt inference engine to various GPU/NPU hardware architectures and ensure high universality.
  • Lead design and optimization of distributed parallel solutions (tensor/pipeline/sequence/MoE) for large models.
  • Stay current with large-model inference tech (GPU HPC, distributed parallelism, cache optimization) and benchmark against frameworks like vLLM and TensorRT-LLM.

Skills

C/C++
Python
CUDA programming
GPU architecture
Deep learning inference
Nsight/Profiler

Education

Bachelor's/Master's in computing

Tools

Nsight
Profiler
TensorRT

Job description

Team Introduction

Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation/advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It provides powerful Machine Learning computing power for internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses.

We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.

Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.

Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to ByteDance and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

Responsibilities
  • Responsible for the iteration of the underlying architecture of the large model inference engine and end-to-end GPU performance optimization, through means such as operator fusion and compilation optimization, deeply optimizing GPU memory access, computing pipeline, and Stream asynchronous scheduling, eliminating inference computing bottlenecks, improving single-card inference throughput, and reducing inference latency.
  • Adapt to all series of GPU/NPU hardware architectures, refine the universality of the inference engine and hardware adaptability, and build a high-performance, low-loss underlying base for large model inference.
  • Lead the design, development, and optimization of distributed parallel solutions for large model inference scenarios, with a focus on implementing multi-dimensional parallel strategies such as tensor parallelism (TP), pipeline parallelism (PP), sequence parallelism, and MoE expert parallelism, to address core issues such as multi-card splitting and deployment of ultra-large models, high cross-card communication overhead, load imbalance, and low parallel efficiency.
  • Follow up on cutting-edge technologies such as global large model inference, GPU high-performance computing, distributed parallelism, and cache optimization, benchmark against mainstream inference frameworks such as vLLM and TensorRT-LLM, complete the implementation of solutions and technological innovation, continuously iterate and optimize the performance and cost advantages of the inference system, and build the core technological barriers of the team.
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's/ Master's degree in computing or a related discipline.
  • Solid foundation in computer low-level knowledge, proficient in C/C++ and Python programming, skilled in CUDA programming and familiar with GPU hardware architecture principles, and well-versed in GPU memory models, computing scheduling, and communication mechanisms;
  • Proficiently master the underlying development and implementation of various basic operators in Deep learning, be well-versed in GPU adaptation and optimization of core operators such as matrix operations, normalization, and activation functions, and be able to independently complete operator handwritten reconstruction, memory access optimization, vectorization acceleration, and precision alignment to ensure high performance and high stability of operator inference.
  • Familiar with the end-to-end process of deep learning inference compilation, understand core compilation technologies such as computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation, and be able to simplify the inference process, reduce GPU memory usage, and decrease inference latency through compilation-level improvements, thereby significantly enhancing the throughput efficiency of model inference.
  • Proficient in using GPU performance analysis tools such as Nsight and Profiler, able to accurately identify performance bottlenecks such as computing power waste, memory access blockage, and scheduling redundancy during the inference process, possess the thinking of software-hardware collaborative optimization, capable of outputting systematic optimization solutions and completing implementation iterations, and adaptable to the requirements of industrial-level high-concurrency, low-latency inference business.
  • Possess good cross-team collaboration skills, communication and presentation skills, and document writing skills, have strong sense of responsibility and stress tolerance, and be able to drive the resolution of complex technical issues and the implementation of projects.
Preferred Qualifications
  • Thoroughly understand the core principles of large model i
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Machine Learning Backend Engineer Graduate (AML MLdev) - 2027 Start
Machine Learning Backend Engineer Graduate (AML MLdev) - 2027 Start

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start Technology - Backend Bachelor/Master Graduate - 2027 Start Singapore Regular
Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start Technology - Backend Bachelor/Master Graduate - 2027 Start Singapore Regular

Bytedance • Singapore

On-site
SGD 180,000 - 280,000
Backend Inference Engineer — High-Performance GPU
Backend Inference Engineer — High-Performance GPU

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 100,000 - 160,000
Applied Machine Learning Orchestration Graduate (AML) - 2027 Start
Applied Machine Learning Orchestration Graduate (AML) - 2027 Start

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 70,000 - 90,000
Meals provided
Research Scientist - Large-Scale Machine Learning Systems (SysML) - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - Large-Scale Machine Learning Systems (SysML) - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Stock options
Positive team atmosphere
Career growth opportunity
+5
Graduate ML System Engineer - Large-Scale Systems
Graduate ML System Engineer - Large-Scale Systems

ByteDance • Singapore

On-site
SGD 60,000 - 80,000
Graduate Backend Inference Framework Engineer - Scalable AI
Graduate Backend Inference Framework Engineer - Scalable AI

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Production Engineer Intern (AML Serving) - 2027 Start
Production Engineer Intern (AML Serving) - 2027 Start

ByteDance • Singapore

On-site
SGD 20,000 - 29,000