Graduate Software Engineer - Cloud-Native GPU Inference

ByteDance

San Jose (CA)

On-site

USD 162,000 - 317,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401k with match
Paid parental leave
Disability Insurance
Life Insurance
Wellbeing Benefits
Paid Holidays
Paid Sick Days
Paid Personal Time

Job summary

ByteDance's Inference Infrastructure team is seeking engineers to design, build, and operate cloud-native GPU-accelerated ML infrastructure at scale, including vLLM, SGLang, and TensorRT-LLM work.

You will join a world-class team within Core Compute Infrastructure, contributing to open-source ecosystems, scheduling, and orchestration across multi-cloud environments. A PhD and deep expertise in distributed systems help you drive high-performance inference at scale.

Qualifications

  • PhD in Software Development, Computer Science, Computer Engineering, or related discipline.
  • Experience with large model inference and distributed systems.
  • Hands-on experience building cloud or ML infrastructure incl. resource management, scheduling, monitoring.
  • Solid knowledge of container/orchestration tech (Docker, Kubernetes).
  • Proficiency in Go, Rust, Python, or C++.

Responsibilities

  • Design and build large-scale container-based cluster management and orchestration systems.
  • Architect next-generation cloud-native GPU and AI accelerator infrastructure.
  • Collaborate to deliver inference solutions using vLLM, SGLang, TensorRT-LLM.
  • Stay current with open source and ML infra advances; implement best practices.
  • Write high-quality, production-ready code.

Skills

Go
Rust
Python
C++
Docker
Kubernetes
Distributed systems
Cloud infrastructure

Education

PhD in Software Development, Computer Science, Computer Engineering

Tools

Docker
Kubernetes

Job description

ByteDance's Inference Infrastructure team is seeking engineers to design, build, and operate cloud-native GPU-accelerated ML infrastructure at scale, including vLLM, SGLang, and TensorRT-LLM work.

You will join a world-class team within Core Compute Infrastructure, contributing to open-source ecosystems, scheduling, and orchestration across multi-cloud environments. A PhD and deep expertise in distributed systems help you drive high-performance inference at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Graduate Cloud-Native AI Inference Engineer
Graduate Cloud-Native AI Inference Engineer

ByteDance • Seattle (WA)

On-site
USD 154,000 - 301,000
Graduate Cloud-Native LLM Inference Engineer
Graduate Cloud-Native LLM Inference Engineer

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
AI Compute Systems Engineer (PhD Grad) 2027 Start
AI Compute Systems Engineer (PhD Grad) 2027 Start

Pangle • San Jose (CA)

Hybrid
USD 162,000 - 317,000
AI Systems & Cloud Infrastructure Intern
AI Systems & Cloud Infrastructure Intern

ByteDance • San Jose (CA)

On-site
USD 50,000 - 83,000
Health insurance
Life insurance
Wellbeing benefits
+3
Graduate Software Engineer, AI Compute & Cloud Systems
Graduate Software Engineer, AI Compute & Cloud Systems

ByteDance • Seattle (WA)

On-site
USD 154,000 - 301,000
Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Graduate Research Engineer, AI Infra & Compute
Graduate Research Engineer, AI Infra & Compute

ByteDance • Seattle (WA)

On-site
USD 130,000 - 180,000