Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance

San Jose (CA)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance's Volcano Ark team is seeking a software engineer to design and develop resource scheduling systems for machine learning workloads across data centers and clusters. The role involves optimizing orchestration of GPUs, CPUs, storage, and networking to support offline training and online inference.

The candidate should have a CS degree, strong programming skills (Go/Java/Python), experience with ML frameworks (TensorFlow/PyTorch), and solid knowledge of Kubernetes, Docker, and distributed

Qualifications

  • Bachelor's or Master's degree in Computer Science or a related discipline.
  • Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python.
  • Solid foundation in algorithms, data structures, and good coding habits.
  • Familiar with at least one mainstream ML framework (TensorFlow, PyTorch).
  • Familiar with Kubernetes architecture and container tech (Docker, container, Kata).
  • Understands distributed systems and has worked on large-scale distributed systems.

Responsibilities

  • Design and develop resource scheduling systems for ML workloads across Volcano Ark and ML platform products.
  • Optimize orchestration and scheduling of GPUs, CPUs, storage, and network resources across data centers and clusters.
  • Support offline training, online inference, and other workloads with multi-tenant isolation to improve utilization and efficiency.

Skills

Go
Java
Python
Linux
Distributed systems
Kubernetes
Docker
TensorFlow
PyTorch

Education

Bachelor's degree in Computer Science or related

Tools

Docker
Kubernetes

Job description

Join us as we work together to inspire creativity and enrich life around the globe.

Location:

San Jose

Team:

Technology

Employment Type:

Regular

Job Code:

A49989

Responsibilities

Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform in China's large model market by product capability and market share. The platform provides end-to-end services including model inference, evaluation, fine-tuning, AI application development, and a plugin ecosystem. Volcano Ark hosts Doubao and leading industry large models, and supports enterprise AI adoption through stable, secure, and trusted solutions as well as professional algorithm and technical services. Data AML is ByteDance's machine learning platform team. It provides training and inference systems for recommendation, advertising, computer vision, speech, and NLP scenarios across products such as Douyin, Toutiao, and Xigua Video. The team also supports internal business teams with large-scale machine learning compute, explores general and innovative algorithms for business problems, and offers core machine learning and recommendation system capabilities to external enterprise customers through Volcano Engine. We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

  • Design and develop resource scheduling systems for machine learning workloads, supporting Volcano Ark and machine learning platform products.
  • Optimize orchestration and scheduling of heterogeneous compute resources, including GPUs, CPUs, and other accelerators, as well as storage resources such as cloud storage and networking resources such as VPC and RDMA, across multiple data centers and clusters.
  • Support scheduling requirements for offline training, online inference, and other workloads under strict multi-tenant isolation, improving overall resource utilization and efficiency.
Qualifications
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science or a related discipline.
  • Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python.
  • Solid foundation in computer science and programming, familiarity with common algorithms and data structures, and good coding habits.
  • Familiar with at least one mainstream machine learning framework, such as TensorFlow, PyTorch, or an internally developed framework.
  • Familiar with Kubernetes architecture and ecosystem, as well as container technologies such as Docker, container, and Kata.
  • Understands distributed system principles and has participated in the design, development, or maintenance of large-scale distributed systems.
Preferred Qualifications
  • Practical experience in large-scale cluster online/offline resource scheduling; source-level understanding of one or more open-source schedulers such as Kubernetes, Volcano, YARN, or Mesos; familiarity with containerization and lightweight virtualization technologies.
  • Deep understanding and practical experience in scheduling topics such as multi-tenant quota governance, preemption, elasticity, fragmentation, tidal scheduling, co-location, and QoS; strong analytical and modeling ability for complex problems; GPU scheduling experience is preferred.
  • Experience in at least one of the following areas: CUDA, RDMA, AI infrastructure, hardware and software co-design, high-performance computing, machine learning hardware architecture such as GPUs, accelerators and networking, ML for systems, or distributed storage.
  • Hands-on experience in cloud-native machine learning systems is preferred.
Job Information
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives.

Our platform connects people from across the globe and so does our workplace.

At ByteDance, our mission is to inspire creativity and enrich life.

To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach.

We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Large Language Model Training System Engineer Graduate (Applied Machine Learning) - 2027 Start [...]
Large Language Model Training System Engineer Graduate (Applied Machine Learning) - 2027 Start [...]

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
LLM Backend Engineer Graduate (Applied Machine Learning) - 2027 Start
LLM Backend Engineer Graduate (Applied Machine Learning) - 2027 Start

Bytedance • San Jose (CA)

On-site
USD 120,000 - 170,000
LLM Backend Engineer Graduate (Applied Machine Learning) - 2027 Start
LLM Backend Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical, dental, and vision insurance
401(k) with company match
Parental leave
+6
Large Language Model Training System Engineer Graduate (Applied Machine Learning) - 2027 Start
Large Language Model Training System Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+6
Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD) PhD Graduates - 2027 Start San Jose Regular
Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD) PhD Graduates - 2027 Start San Jose Regular

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start
Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+4
Software Engineer Intern (AML-Engine-Orchestration) - 2027 Start Undergraduate/Master Intern - [...]
Software Engineer Intern (AML-Engine-Orchestration) - 2027 Start Undergraduate/Master Intern - [...]

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 20,000 - 27,000
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD) PhD Graduates - 2027 Start San Jose Regular
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD) PhD Graduates - 2027 Start San Jose Regular

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 162,000 - 317,000
Health insurance
401(k) with company match
Parental leave
+6
Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)
Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000