Graduate Backend Inference Engine Engineer

ByteDance

San Jose (CA)

On-site

USD 128,000 - 256,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(k) with company match
Paid parental leave
Disability coverage
Life insurance
Wellbeing benefits
10 paid holidays per year
10 paid sick days per year
Paid Personal Time

Job summary

ByteDance in San Jose seeks a Senior GPU Inference Engineer to drive the iteration of large model inference engines and optimize GPU memory, latency, and throughput across multi-card systems.

You will design distributed parallel solutions (TP/PP/sequence/MoE) and push performance across vLLM, TensorRT-LLM, and related frameworks while collaborating with cross-functional teams to scale cutting-edge AI capabilities for internal products.

Qualifications

  • Bachelor's/Master's degree in Software Development, CS, or related technical discipline.
  • Solid foundation in computer low-level knowledge; proficient in C/C++ and Python.
  • Proficient in CUDA programming; familiar with GPU memory models and scheduling.
  • Experience with deep learning operators, graph optimization, and memory optimization.
  • Experience using GPU performance tools like Nsight/Profiler and profiling for bottlenecks.
  • Familiar with large model inference frameworks and multi-card parallelism concepts.

Responsibilities

  • Iterate architecture of large model inference engine and optimize GPU latency and throughput.
  • Adapt to GPU/NPU hardware architectures and ensure high performance across devices.
  • Lead design and optimization of distributed parallel solutions (TP/PP/sequence/MoE).
  • Stay updated on global large model inference tech and benchmark against frameworks like vLLM and TensorRT-LLM.

Skills

C/C++
Python
CUDA
GPU hardware
Performance analysis

Education

Bachelor's/Master's degree in CS/Engineering

Tools

Nsight
Profiler
TensorRT-LLM

Job description

ByteDance in San Jose seeks a Senior GPU Inference Engineer to drive the iteration of large model inference engines and optimize GPU memory, latency, and throughput across multi-card systems.

You will design distributed parallel solutions (TP/PP/sequence/MoE) and push performance across vLLM, TensorRT-LLM, and related frameworks while collaborating with cross-functional teams to scale cutting-edge AI capabilities for internal products.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Large-Model Inference Runtime Engineer
Large-Model Inference Runtime Engineer

Bytedance • San Jose (CA)

On-site
USD 128,000 - 256,000
LLM/VLM Inference Optimization Research Engineer
LLM/VLM Inference Optimization Research Engineer

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+1
Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Graduate GPU AI Platform Engineer — System Optimization
Graduate GPU AI Platform Engineer — System Optimization

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

Bytedance • San Jose (CA)

On-site
USD 128,000 - 256,000
AI Infrastructure & ML Systems Intern
AI Infrastructure & ML Systems Intern

ByteDance • San Jose (CA)

On-site
USD 97,000 - 138,000
Health insurance
Housing allowance
Paid holidays
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
LLM Training & Inference Scientist (GPU-Optimized)
LLM Training & Inference Scientist (GPU-Optimized)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3