Senior Machine Learning Engineer - Orchestration

ByteDance

Seattle (WA)

On-site

USD 207,000 - 368,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ByteDance in Seattle seeks a senior distributed ML systems engineer to optimize scheduling, build scalable training/inference runtimes, and drive online orchestration for next-gen recommender models. You will work across Kubernetes-based frameworks and large-scale ML workloads, contributing to both research and production platforms.

The role requires 5+ years in Go or Python, strong distributed systems knowledge, and ability to design robust, scalable engineering solutions in a fast-paced team

Qualifications

  • Bachelor's degree or above in Computer Science or similar field.
  • At least 5 years of experience with Go or Python in a Linux environment.
  • Familiar with distributed scheduling frameworks and ML systems.
  • Master the principles of distributed systems and design/maintain large-scale systems.
  • Strong analytical, documentation, and self-motivation skills.

Responsibilities

  • Optimize resource efficiency in distributed orchestration and scheduling.
  • Develop scheduling frameworks around Kubernetes/Godel ecosystem for various scenarios.
  • Extend AutoScaling and automatic parallelization for models and operations.
  • Manage preemption/eviction, cross-cluster resource docking, and multi-datacenter runtime.
  • Build training system architecture for ultra-large recommendation models.
  • Design distributed training runtimes and ML model synchronization.
  • Interface with platform to improve diagnosability of distributed training.
  • Construct online orchestration for next-gen Recommender system and inference.

Skills

Go
Python
Distributed systems knowledge
Strong coding skills

Education

Bachelor's degree in Computer Science or similar

Tools

Kubernetes
Yarn
Flink
MapReduce
Mesos
Celery

Job description

Location:

Seattle

Team:

Technology

Employment Type:

Regular

Job Code:

A230840

Share this listing:

About the Team:

Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation, advertising, CV, speech, and NLP for businesses such as Douyin. It provides powerful Machine Learning computing power to internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses. At the same time, it also provides some core capabilities of Machine Learning and Recommender systems to external enterprise customers through Volcano Engine. In addition, AML also conducts some cutting-edge research in fields such as Al for Science and scientific computing.

Responsibilities
  • Optimizing resource efficiency in distributed orchestration and scheduling, through engineering means, enhances the scale of business/models supported per unit of computing power:
  • Use/secondarily develop distributed scheduling frameworks around the Kubernetes/Godel ecosystem, make reasonable selections in different business scenarios, and optimize scheduling strategies for cluster utilization/uniformity based on the characteristics of different scenarios;
  • Connect/extend AutoScaling for various models and business operations, as well as automatic parallelization tasks. Through the method of load modeling and analysis of different models, automatically optimize resource requests for models, optimize resource utilization efficiency at scale, and achieve global optimality;
  • Responsible for the preemption/eviction function of services with different priorities; responsible for the borrowing/mixed deployment docking work among different types of resources in different clusters; responsible for the scheduling/load adaptation in scenarios of multiple data centers, multiple regions, and multiple clouds;
  • Build a training system architecture for next-generation ultra-large and ultra-deep recommendation models:
  • Build a flexible and robust distributed training runtime around ultra-large-scale embedding and ultra-large-scale GPU synchronization training;
  • Design and optimize distributed computing APis and runtime for future-oriented research paradigms of recommended advertising models (e.g., RL/finetune/distillation);
  • Interface with the platform to optimize the diagnosability and usability of distributed training.
  • Construct an online orchestration architecture for the next-generation Recommender system:
  • Build a robust and stable distributed model inference architecture around the online training scenario of ultra-large-scale embeddings;
  • Optimize the usability of the online architecture of the recommended advertising model and the MLops process by integrating the research and experimental model of the business.
Qualifications
Minimum Qualifications
  • Bachelor's degree or above in Computer Science or similar field of study.
  • At least 5 years of experience with proficiency in at least one programming language among Go/Python in a Linux environment, with excellent hands-on coding skills.
  • Familiar with some open-source distributed scheduling frameworks, such as Kubernetes (K8S), Yarn (as well as the Big data frameworks Flink and MapReduce in the Hadoop ecosystem), Mesos, Celery, and has rich practical and development experience in Machine Learning systems.
  • Master the principles of distributed systems, and have participated in the design, development, and maintenance of large-scale distributed systems.
  • Have excellent logical analysis skills, capable of reasonably abstracting and splitting business logic.
  • Have a strong sense of work responsibility, good learning ability, communication skills, and self-motivation, and be able to respond and act quickly.
  • Have good work documentation habits, and timely write and update work processes and Technical Documentation as required.
Preferred Qualifications
  • Familiar with at least one mainstream Machine Learning framework (PyTorch / TensorFlow);
  • Have experience in one of the following areas: Al Infrastructure, HW/SW Co-Design, High Performance Computing, ML Hardware Architecture (GPU, Accelerators, Networking);
  • Some experience in using/designing open-source training orchestration systems, such as veRL, VLLM, Ray,TFX. Those with development experience in at least one of them are preferred.
Job Information

The base salary range for this position in the selected city is $207480 - $368220 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:
  • 1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  • 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and
  • 3. Exercising sound judgment.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)
Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Senior Software Engineer, AI Infrastructure-Compute Efficiency
Senior Software Engineer, AI Infrastructure-Compute Efficiency

ByteDance • Seattle (WA)

On-site
USD 207,000 - 368,000
Machine Learning Backend Engineer Graduate (AML MLDev) - 2027 Start
Machine Learning Backend Engineer Graduate (AML MLDev) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+9
Machine Learning Engineer Intern (AML-Engine-Orchestration) - 2027 Start
Machine Learning Engineer Intern (AML-Engine-Orchestration) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 51,000 - 73,000
Housing allowance
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000
Software Engineer Intern (AML-Engine-Orchestration) - 2027 Start
Software Engineer Intern (AML-Engine-Orchestration) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 34,000 - 55,000
Software Engineer Intern (AML-Engine-Orchestration) - 2027 Start
Software Engineer Intern (AML-Engine-Orchestration) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 51,000 - 73,000
Software Engineer Graduate (AML-Engine-Forge Platform) - 2027 Start
Software Engineer Graduate (AML-Engine-Forge Platform) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Health insurance
401(k) matching
Parental leave
+6
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD) PhD Graduates - 2027 Start San Jose Regular
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD) PhD Graduates - 2027 Start San Jose Regular

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 162,000 - 317,000
Health insurance
401(k) with company match
Parental leave
+6