ML Scheduler Engineer: Optimizing GPU/CPU Orchestration

ByteDance

San Jose (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ByteDance's Volcano Ark team is seeking a software engineer to design and develop resource scheduling systems for machine learning workloads across data centers and clusters. The role involves optimizing orchestration of GPUs, CPUs, storage, and networking to support offline training and online inference.

The candidate should have a CS degree, strong programming skills (Go/Java/Python), experience with ML frameworks (TensorFlow/PyTorch), and solid knowledge of Kubernetes, Docker, and distributed

Qualifications

  • Bachelor's or Master's degree in Computer Science or a related discipline.
  • Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python.
  • Solid foundation in algorithms, data structures, and good coding habits.
  • Familiar with at least one mainstream ML framework (TensorFlow, PyTorch).
  • Familiar with Kubernetes architecture and container tech (Docker, container, Kata).
  • Understands distributed systems and has worked on large-scale distributed systems.

Responsibilities

  • Design and develop resource scheduling systems for ML workloads across Volcano Ark and ML platform products.
  • Optimize orchestration and scheduling of GPUs, CPUs, storage, and network resources across data centers and clusters.
  • Support offline training, online inference, and other workloads with multi-tenant isolation to improve utilization and efficiency.

Skills

Go
Java
Python
Linux
Distributed systems
Kubernetes
Docker
TensorFlow
PyTorch

Education

Bachelor's degree in Computer Science or related

Tools

Docker
Kubernetes

Job description

ByteDance's Volcano Ark team is seeking a software engineer to design and develop resource scheduling systems for machine learning workloads across data centers and clusters. The role involves optimizing orchestration of GPUs, CPUs, storage, and networking to support offline training and online inference.

The candidate should have a CS degree, strong programming skills (Go/Java/Python), experience with ML frameworks (TensorFlow/PyTorch), and solid knowledge of Kubernetes, Docker, and distributed

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Compute Orchestration & ML Scheduling
Senior Software Engineer, Compute Orchestration & ML Scheduling

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Medical insurance
401(k) plan with company match
Parental leave
+6
ML Platform & Orchestration Engineer
ML Platform & Orchestration Engineer

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
ML Orchestration Engineer for Scalable Model Serving
ML Orchestration Engineer for Scalable Model Serving

Bytedance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical insurance
401(k) match
Paid parental leave
+2
AI Scheduling Systems Engineer (Graduate)
AI Scheduling Systems Engineer (Graduate)

ByteDance • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical Insurance
Dental Insurance
Vision Insurance
+9
ML Systems Engineer: Kubernetes Orchestration & Model Serving
ML Systems Engineer: Kubernetes Orchestration & Model Serving

ByteDance • Seattle (WA)

On-site
USD 140,000 - 210,000
Senior ML Platform Engineer: Distributed Training & Scheduling
Senior ML Platform Engineer: Distributed Training & Scheduling

ByteDance • Seattle (WA)

On-site
USD 207,000 - 368,000
Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start
Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 120,000 - 180,000
Founding ML Platforms Engineer — GPU Orchestrator
Founding ML Platforms Engineer — GPU Orchestrator

cumulus labs • San Francisco (CA)

On-site
USD 140,000 - 180,000
ML Infra Engineer Intern: Build Scalable Orchestration
ML Infra Engineer Intern: Build Scalable Orchestration

ByteDance • Seattle (WA)

On-site
USD 34,000 - 55,000
Senior AI Scheduling & GPU Orchestration Engineer
Senior AI Scheduling & GPU Orchestration Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000