Graduate ML Engineer: Scalable AI Infrastructure

ByteDance

San Jose (CA)

On-site

USD 162,000 - 317,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
Short-term and long-term disability
Life insurance
Wellbeing benefits
10 paid holidays per year
10 paid sick days per year
17 days Paid Personal Time

Job summary

ByteDance is seeking highly capable researchers to join the Data-AML-Engine Orchestration team in California. You will help design large-scale ML infrastructure powering model serving across ByteDance products including TikTok, focusing on orchestration, scheduling, and resource management for diverse compute environments.

Successful candidates will work on systems affecting GPU utilization, latency, reliability, and productivity.

Qualifications

  • PhD in Computer Science, Software Engineering, AI or related field.
  • Proficiency in Go, C++, or Python with solid data structures and algorithms.
  • Familiarity with Linux, OS concepts, networks, concurrency and distributed systems.
  • Hands-on, exploratory, with ability to analyze systems via metrics and logs.
  • Systematic, quantitative problem solving with measurable improvements.
  • Demonstrated ownership through coursework, research, internships, or open‑source work.

Responsibilities

  • Design & build foundational orchestration capabilities for ML platforms (Kubernetes Operators, containers, lifecycle management).
  • Create multi-tenant resource and quota systems, improve GPU utilization, cost efficiency via resource pooling.
  • Build lifecycle orchestration for online model serving: deployment, upgrades, rollback, autoscaling, multi-cluster ops, DR.
  • Develop serving orchestration and traffic management for disaggregated clusters: topology-aware scheduling, KV Cache affinity, QoS/SLA.

Skills

PhD in CS
Go/C++/Python
Linux & distributed systems
Problem solving
Ownership & collaboration

Education

PhD in Computer Science

Tools

Kubernetes
Container runtimes
Volcano/OpenKruise

Job description

ByteDance is seeking highly capable researchers to join the Data-AML-Engine Orchestration team in California. You will help design large-scale ML infrastructure powering model serving across ByteDance products including TikTok, focusing on orchestration, scheduling, and resource management for diverse compute environments.

Successful candidates will work on systems affecting GPU utilization, latency, reliability, and productivity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Graduate ML Engineer: Orchestration & Online Model Serving
Graduate ML Engineer: Orchestration & Online Model Serving

ByteDance • San Jose (CA)

On-site
USD 180,000 - 230,000
Graduate ML Platform Engineer - Orchestration & Serving
Graduate ML Platform Engineer - Orchestration & Serving

ByteDance • San Jose (CA)

On-site
USD 180,000 - 240,000
Graduate Software Engineer: ML Infra & Orchestration
Graduate Software Engineer: ML Infra & Orchestration

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000
Graduate ML Systems Engineer - AML Orchestration
Graduate ML Systems Engineer - AML Orchestration

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 162,000 - 317,000
Health insurance
401(k) with company match
Parental leave
+6
ML Platform Engineer Intern
ML Platform Engineer Intern

ByteDance • San Jose (CA)

On-site
USD 51,000 - 73,000
Housing allowance
ML Platform Orchestration Intern
ML Platform Orchestration Intern

ByteDance • San Jose (CA)

On-site
USD 51,000 - 73,000
Graduate Research Scientist, AI Systems & Infrastructure
Graduate Research Scientist, AI Systems & Infrastructure

ByteDance • San Jose (CA)

On-site
USD 218,000 - 388,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+6
Research Scientist: AI Infrastructure & Large-Scale ML
Research Scientist: AI Infrastructure & Large-Scale ML

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
ML Platform Engineer - Orchestration & GPU-Driven Serving
ML Platform Engineer - Orchestration & GPU-Driven Serving

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
ML Platform Intern: Build Scalable Serving & Orchestration
ML Platform Intern: Build Scalable Serving & Orchestration

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 20,000 - 33,000