Senior ML Platform Engineer: Distributed Training & Scheduling

ByteDance

Seattle (WA)

On-site

USD 207,000 - 368,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ByteDance in Seattle seeks a senior distributed ML systems engineer to optimize scheduling, build scalable training/inference runtimes, and drive online orchestration for next-gen recommender models. You will work across Kubernetes-based frameworks and large-scale ML workloads, contributing to both research and production platforms.

The role requires 5+ years in Go or Python, strong distributed systems knowledge, and ability to design robust, scalable engineering solutions in a fast-paced team

Qualifications

  • Bachelor's degree or above in Computer Science or similar field.
  • At least 5 years of experience with Go or Python in a Linux environment.
  • Familiar with distributed scheduling frameworks and ML systems.
  • Master the principles of distributed systems and design/maintain large-scale systems.
  • Strong analytical, documentation, and self-motivation skills.

Responsibilities

  • Optimize resource efficiency in distributed orchestration and scheduling.
  • Develop scheduling frameworks around Kubernetes/Godel ecosystem for various scenarios.
  • Extend AutoScaling and automatic parallelization for models and operations.
  • Manage preemption/eviction, cross-cluster resource docking, and multi-datacenter runtime.
  • Build training system architecture for ultra-large recommendation models.
  • Design distributed training runtimes and ML model synchronization.
  • Interface with platform to improve diagnosability of distributed training.
  • Construct online orchestration for next-gen Recommender system and inference.

Skills

Go
Python
Distributed systems knowledge
Strong coding skills

Education

Bachelor's degree in Computer Science or similar

Tools

Kubernetes
Yarn
Flink
MapReduce
Mesos
Celery

Job description

ByteDance in Seattle seeks a senior distributed ML systems engineer to optimize scheduling, build scalable training/inference runtimes, and drive online orchestration for next-gen recommender models. You will work across Kubernetes-based frameworks and large-scale ML workloads, contributing to both research and production platforms.

The role requires 5+ years in Go or Python, strong distributed systems knowledge, and ability to design robust, scalable engineering solutions in a fast-paced team

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Graduate Software Engineer: ML Infra & Orchestration
Graduate Software Engineer: ML Infra & Orchestration

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000
ML Systems Engineer: Kubernetes Orchestration & Model Serving
ML Systems Engineer: Kubernetes Orchestration & Model Serving

ByteDance • Seattle (WA)

On-site
USD 140,000 - 210,000
ML Platform & Orchestration Engineer
ML Platform & Orchestration Engineer

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Software Engineer, Compute Orchestration & ML Scheduling
Senior Software Engineer, Compute Orchestration & ML Scheduling

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Medical insurance
401(k) plan with company match
Parental leave
+6
Tech Lead — Scalable ML Platform Engineer
Tech Lead — Scalable ML Platform Engineer

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+2
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Senior Physical AI Infra Engineer — Scalable ML Infra
Senior Physical AI Infra Engineer — Scalable ML Infra

Bytedance • Seattle (WA)

On-site
USD 207,000 - 368,000
Medical insurance
401(k) with company match
Paid parental leave
+3
ML Orchestration Engineer for Scalable Model Serving
ML Orchestration Engineer for Scalable Model Serving

Bytedance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical insurance
401(k) match
Paid parental leave
+2
Research Scientist, AI Infrastructure & Distributed ML
Research Scientist, AI Infrastructure & Distributed ML

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior AI Infra Engineer: Scale Cloud & ML Pipelines
Senior AI Infra Engineer: Scale Cloud & ML Pipelines

ByteDance • Seattle (WA)

On-site
USD 140,000 - 210,000