Large-Model Inference Runtime Engineer

Bytedance

San Jose (CA)

On-site

USD 128,000 - 256,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a highly skilled software engineer to advance its data AML mid-platform, focusing on large model inference and GPU optimization. The role emphasizes architecture iteration, end-to-end performance tuning, and cross-team collaboration across the Technology group in a fast-paced environment.

You will work on cutting-edge inference systems powering Douyin, Jinri Toutiao, and Xigua Video, with opportunities for growth and impact.

Qualifications

  • Bachelor's or Master's degree in Software Development, Computer Science, Computer Engineering, or related technical discipline.
  • Solid foundation in low-level computer knowledge; proficient in C/C++ and Python; CUDA programming; familiar with GPU hardware and memory models.
  • Proficient in developing and optimizing deep learning operators; independent in operator reconstruction, memory access optimization, vectorization, and precision alignment.

Responsibilities

  • Iterate the architecture of the large model inference engine and optimize GPU performance (fusion, compilation, memory, scheduling).
  • Adapt to GPU/NPU hardware architectures; refine universality of the inference engine and hardware adaptability.
  • Design, develop, and optimize distributed parallel solutions for large model inference (tensor/pipeline/sequence/MoE parallelism).
  • Follow cutting-edge large model inference tech, benchmark against frameworks like vLLM and TensorRT-LLM; iterate to improve performance and cost.

Skills

C/C++
Python
CUDA programming
GPU architecture
Deep learning operators
GPU memory models
Nsight/Profiler
Cross-team collaboration
Performance optimization

Education

Bachelor's/Master's in CS/Engineering

Tools

Nsight
Profiler

Job description

ByteDance is seeking a highly skilled software engineer to advance its data AML mid-platform, focusing on large model inference and GPU optimization. The role emphasizes architecture iteration, end-to-end performance tuning, and cross-team collaboration across the Technology group in a fast-paced environment.

You will work on cutting-edge inference systems powering Douyin, Jinri Toutiao, and Xigua Video, with opportunities for growth and impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
High-Performance ML Backend Engineer
High-Performance ML Backend Engineer

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Research Scientist, AI Infrastructure & Distributed ML
Research Scientist, AI Infrastructure & Distributed ML

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Graduate GPU AI Platform Engineer — System Optimization
Graduate GPU AI Platform Engineer — System Optimization

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Graduate AI Model Optimization Engineer
Graduate AI Model Optimization Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

Bytedance • San Jose (CA)

On-site
USD 128,000 - 256,000
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Graduate ML Engineer: Scalable AI Infrastructure
Graduate ML Engineer: Scalable AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6