LLM/VLM Inference Optimization Research Engineer

Bytedance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
Holiday and paid time off

Job summary

ByteDance Seed Infra in San Jose is seeking a Research Engineer to design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs. You will work on inference engines, serving frameworks, and end-to-end deployment pipelines, aiming to reduce latency and increase throughput.

Ideal candidates will have strong C/C++ and Python skills, hands-on ML framework experience (PyTorch/TensorFlow), and background in GPU-accelerated optimization.

Qualifications

  • Bachelor's degree or above in Computer Science, Electrical Engineering, Software Engineering, or a related field.
  • Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming; familiarity with containerization and server-side debugging.
  • Hands-on experience with at least one mainstream machine learning framework (e.g., PyTorch, TensorFlow).
  • Experience deploying or optimizing LLM/VLM inference at production scale, with demonstrated impact on latency, throughput, or serving cost.
  • Familiarity with GPU architecture and experience optimizing compute-intensive operators (e.g., FlashAttention, GEMM, GEMV, Conv2D).

Responsibilities

  • Design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs, covering inference engines, serving frameworks, and end-to-end deployment pipelines.
  • Build state-of-the-art model inference engines through advanced performance optimization techniques such as compiler-level optimizations, parallel computing, graph fusion, efficient CUDA kernel development, low-precision computation, streaming inference, speculative decoding, and high-concurrency request optimization.
  • Collaborate with other research teams to identify performance bottlenecks, conduct performance analysis, and optimize large models; contribute to model toolchains and the broader technical ecosystem.

Skills

C/C++
Python
Algorithms
Data structures
Systems programming
Containerization
PyTorch/TensorFlow
CUDA
GPU architectures

Education

Bachelor's degree or above in CS/EE/SE

Tools

TensorRT
Triton
CUTLASS
CUDA/OpenCL

Job description

ByteDance Seed Infra in San Jose is seeking a Research Engineer to design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs. You will work on inference engines, serving frameworks, and end-to-end deployment pipelines, aiming to reduce latency and increase throughput.

Ideal candidates will have strong C/C++ and Python skills, hands-on ML framework experience (PyTorch/TensorFlow), and background in GPU-accelerated optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM/VLM Inference Optimization Engineer
LLM/VLM Inference Optimization Engineer

ByteDance • Seattle (WA)

On-site
USD 232,000 - 428,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
LLM Training & Inference Scientist (GPU-Optimized)
LLM Training & Inference Scientist (GPU-Optimized)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Graduate Research Scientist, LLM & AI Infrastructure
Graduate Research Scientist, LLM & AI Infrastructure

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
LLM Storage Systems Research Engineer
LLM Storage Systems Research Engineer

ByteDance • San Jose (CA)

On-site
USD 156,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior LLM Storage Systems Engineer & Researcher
Senior LLM Storage Systems Engineer & Researcher

ByteDance • Seattle (WA)

On-site
USD 202,160 - 368,220
Medical, dental, vision insurance
401(k) matching
Paid parental leave
+1
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
LLM Research Scientist - Next-Gen Models & Innovation
LLM Research Scientist - Next-Gen Models & Innovation

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental & vision insurance
401(k) with company match
Paid parental leave
+5
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA