Graduate Cloud-Native LLM Inference Engineer

ByteDance

San Jose (CA)

On-site

USD 162,000 - 317,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) with company match
Paid parental leave
Disability coverage
Life insurance
Wellbeing benefits
10 holidays per year
10 sick days per year
Paid Personal Time

Job summary

ByteDance is seeking engineers for its Inference Infrastructure team to advance the scale and efficiency of AI inference across multi-cloud data centers. You will work on Kubernetes-native control planes, GPU-accelerated ML platforms, and open-source tooling to enable production-grade AI workloads.

The role emphasizes collaboration across global teams, contribution to open-source, and tackling ambitious cloud-native challenges in a hyper-scale environment.

Qualifications

  • PhD in Computer Science or related technical discipline.
  • Strong understanding of large model inference and distributed systems.
  • Hands-on experience building cloud or ML infrastructure.
  • Knowledge of container/orchestration technologies (Docker, Kubernetes).
  • Proficiency in at least one language (Go, Rust, Python, or C).

Responsibilities

  • Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next-generation cloud-native GPU and AI accelerator infrastructure for cost-efficient ML platforms.
  • Collaborate across teams to deliver inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • Stay current with open source and AI/ML advances; integrate best practices into production systems.
  • Write high-quality, production-ready code that is maintainable, testable, and scalable.

Skills

Large model inference
Distributed systems
Cloud infrastructure
Docker & Kubernetes
Go / Rust / Python / C
GPU acceleration

Education

PhD in Computer Science or related

Tools

vLLM
SGLang
TensorRT-LLM
Ray

Job description

ByteDance is seeking engineers for its Inference Infrastructure team to advance the scale and efficiency of AI inference across multi-cloud data centers. You will work on Kubernetes-native control planes, GPU-accelerated ML platforms, and open-source tooling to enable production-grade AI workloads.

The role emphasizes collaboration across global teams, contribution to open-source, and tackling ambitious cloud-native challenges in a hyper-scale environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Graduate Cloud-Native AI Inference Engineer
Graduate Cloud-Native AI Inference Engineer

ByteDance • Seattle (WA)

On-site
USD 154,000 - 301,000
Graduate Software Engineer - Cloud-Native GPU Inference
Graduate Software Engineer - Cloud-Native GPU Inference

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
AI Infrastructure Engineer (LLM Training & Inference)
AI Infrastructure Engineer (LLM Training & Inference)

ByteDance • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical Insurance
Dental Insurance
Vision Insurance
+9
Graduate Cloud-Native AI Inference Engineer
Graduate Cloud-Native AI Inference Engineer

ByteDance • Seattle (WA)

On-site
USD 148,000 - 301,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
AI Compute Systems Engineer (PhD Grad) 2027 Start
AI Compute Systems Engineer (PhD Grad) 2027 Start

Pangle • San Jose (CA)

Hybrid
USD 162,000 - 317,000
Graduate Research Engineer, AI Infra & Compute
Graduate Research Engineer, AI Infra & Compute

ByteDance • Seattle (WA)

On-site
USD 130,000 - 180,000
Graduate Software Engineer, AI Compute & Cloud Systems
Graduate Software Engineer, AI Compute & Cloud Systems

ByteDance • Seattle (WA)

On-site
USD 154,000 - 301,000
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
AI Infrastructure Intern — Cloud Compute Platform
AI Infrastructure Intern — Cloud Compute Platform

ByteDance • San Jose (CA)

On-site
USD 51,000 - 73,000
Health insurance from day one
Housing allowance
Paid holidays