ML Infrastructure Engineer Intern

ByteDance

Seattle (WA)

On-site

USD 33,000 - 50,000

Part time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a motivated intern for the Data-AML-Engine Orchestration team in Seattle to help build large-scale ML infrastructure. You will work on orchestration, scheduling, and resource management that connect heterogeneous compute with production ML workloads, influencing GPU utilization, latency, and reliability.

You will explore Kubernetes Operators, multi-tenant scheduling, and online model serving, collaborating with engineers across teams.

Qualifications

  • Pursuing Bachelor's or Master's in CS, SE, AI, or related field.
  • Proficiency in Go, C++, or Python with solid CS fundamentals.
  • Familiar with Linux and OS concepts, networks, concurrency, and distributed systems.
  • Hands-on and exploratory mindset with ability to analyze systems via code, metrics, logs, profiling, and experiments.
  • Systematic and quantitative problem solving with ability to define measurements and validate improvements.
  • Demonstrated ownership through coursework, internships, OSS, or other projects.

Responsibilities

  • Design and build orchestration features for ML platforms, including Kubernetes Operators and lifecycle management for jobs and services.
  • Build multi-tenant resource and quota systems to improve GPU utilization and cost efficiency.
  • Develop serving orchestration for online model deployment, upgrades, autoscaling, and disaster recovery.
  • Contribute to topology-aware scheduling, KV caching, intelligent routing, and QoS/SLA management.

Skills

Go
C++
Python

Education

Bachelor's or Master's in CS/SE/AI or related

Tools

Linux
Kubernetes

Job description

ByteDance is seeking a motivated intern for the Data-AML-Engine Orchestration team in Seattle to help build large-scale ML infrastructure. You will work on orchestration, scheduling, and resource management that connect heterogeneous compute with production ML workloads, influencing GPU utilization, latency, and reliability.

You will explore Kubernetes Operators, multi-tenant scheduling, and online model serving, collaborating with engineers across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer Intern: Build Scalable Orchestration
ML Infra Engineer Intern: Build Scalable Orchestration

ByteDance • Seattle (WA)

On-site
USD 34,000 - 55,000
ML Infra Intern: Build Scalable Orchestration
ML Infra Intern: Build Scalable Orchestration

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 20,000 - 27,000
ML Platform Intern: Build Scalable Serving & Orchestration
ML Platform Intern: Build Scalable Serving & Orchestration

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 20,000 - 33,000
Graduate Software Engineer: ML Infra & Orchestration
Graduate Software Engineer: ML Infra & Orchestration

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000
Graduate ML Orchestration Engineer — Seattle GPU Serving
Graduate ML Orchestration Engineer — Seattle GPU Serving

ByteDance • Seattle (WA)

On-site
USD 140,000 - 190,000
AI Infra Intern: Build Smarter Compute at Scale
AI Infra Intern: Build Smarter Compute at Scale

Bytedance • Seattle (WA), Northern (KY)

Hybrid
USD 34,000 - 55,000
AI Infrastructure Intern: Distributed Training & ML Systems
AI Infrastructure Intern: Distributed Training & ML Systems

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 20,000 - 31,000
ML Systems Engineer: Kubernetes Orchestration & Model Serving
ML Systems Engineer: Kubernetes Orchestration & Model Serving

ByteDance • Seattle (WA)

On-site
USD 140,000 - 210,000
ML Production Engineering Intern - Build Scalable Systems
ML Production Engineering Intern - Build Scalable Systems

ByteDance • San Jose (CA)

On-site
USD 52,000 - 72,000
Health insurance from day one
Wellbeing benefits
Housing allowance for non-remote
Graduate ML Engineer: Scalable AI Infrastructure
Graduate ML Engineer: Scalable AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6