Graduate Cloud-Native AI Inference Engineer

ByteDance

Seattle (WA)

On-site

USD 154,000 - 301,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ByteDance’s Inference Infrastructure team is hiring engineers to design and operate cloud-native, GPU-accelerated ML platforms at scale. You will work on open-source oriented systems for large-scale LLM inference, contributing to scheduling, resource management, and GPU orchestration in a hyper-scale environment.

We seek PhD graduates with strong distributed systems knowledge, experience with Docker/Kubernetes, and proficiency in Go, Rust, Python, or C++.

Qualifications

  • PhD in Software Development, Computer Science, Computer Engineering or related field.
  • Strong understanding of large model inference and distributed/parallel systems.
  • Hands-on experience building cloud or ML infrastructure (resource management, scheduling, monitoring).
  • Solid knowledge of containers and orchestration technologies (Docker, Kubernetes).
  • Proficiency in Go, Rust, Python, or C++.

Responsibilities

  • Design and build large-scale container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms.
  • Collaborate across teams to deliver world-class inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high-quality, production-ready code that is maintainable, testable, and scalable.

Skills

PhD in software development/computer-
Large model inference
Cloud/ML infrastructure
Docker & Kubernetes
Go/Rust/Python/C++

Education

PhD in Computer Science/Engineering or related field

Tools

Docker
Kubernetes
Ray
TensorRT-LLM

Job description

ByteDance’s Inference Infrastructure team is hiring engineers to design and operate cloud-native, GPU-accelerated ML platforms at scale. You will work on open-source oriented systems for large-scale LLM inference, contributing to scheduling, resource management, and GPU orchestration in a hyper-scale environment.

We seek PhD graduates with strong distributed systems knowledge, experience with Docker/Kubernetes, and proficiency in Go, Rust, Python, or C++.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Graduate Cloud-Native LLM Inference Engineer
Graduate Cloud-Native LLM Inference Engineer

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Graduate Software Engineer - Cloud-Native GPU Inference
Graduate Software Engineer - Cloud-Native GPU Inference

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Graduate Software Engineer, AI Compute & Cloud Systems
Graduate Software Engineer, AI Compute & Cloud Systems

ByteDance • Seattle (WA)

On-site
USD 154,000 - 301,000
Graduate Cloud-Native AI Inference Engineer
Graduate Cloud-Native AI Inference Engineer

ByteDance • Seattle (WA)

On-site
USD 148,000 - 301,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
AI Compute Systems Engineer (PhD Grad) 2027 Start
AI Compute Systems Engineer (PhD Grad) 2027 Start

Pangle • San Jose (CA)

Hybrid
USD 162,000 - 317,000
Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Graduate Research Engineer, AI Infra & Compute
Graduate Research Engineer, AI Infra & Compute

ByteDance • Seattle (WA)

On-site
USD 130,000 - 180,000
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Graduate Research Scientist, AI Systems & Infrastructure
Graduate Research Scientist, AI Systems & Infrastructure

ByteDance • San Jose (CA)

On-site
USD 218,000 - 388,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+6
AI Systems & Cloud Infrastructure Intern
AI Systems & Cloud Infrastructure Intern

ByteDance • San Jose (CA)

On-site
USD 50,000 - 83,000
Health insurance
Life insurance
Wellbeing benefits
+3