Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

ByteDance

San Jose (CA)

On-site

USD 156,000 - 387,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
401(k) plan with company match
Parental leave
Disability insurance
Life insurance
Wellbeing benefits
10 paid holidays per year
10 paid sick days per year
Paid Personal Time

Job summary

ByteDance in San Jose seeks a Senior Software Engineer for Compute Infrastructure (Orchestration & Scheduling). You will scale and enhance Kubernetes-based systems, designing a unified scheduler for diverse workloads across global data centers.

You will also build AI-driven scheduling to optimize CPU/GPU, memory and power usage, and drive next-gen ML compute platforms for fast, reliable ML/LLM training and inference.

Qualifications

  • BS/MS in Computer Science, Computer Engineering or related field with 3+ years of experience; PhD with strong publications may qualify
  • Solid understanding of Unix/Linux environments, distributed and parallel systems, or high-performance networking

Responsibilities

  • Engineer hyper-scale cluster management to improve Kubernetes-based platforms
  • Innovate on core scheduling for containers, VMs, AI/ML workloads
  • Develop an intelligent scheduling system using AI to optimize performance and resource use
  • Lead infrastructure for next-gen ML workloads, ML training/inference
  • Deliver high-quality, maintainable code and stay current with open-source and research advancements

Skills

Python
Go
C++
Rust
Java
Unix/Linux
Kubernetes
Docker

Education

Bachelor's or Master’s in CS/CE

Tools

Docker
Kubernetes

Job description

Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

Location: San Jose

Team: Infrastructure

Employment Type: Regular

Job Code: A228603

Responsibilities
  • Engineer hyper-scale cluster management: Enhance Kubernetes-based cluster platforms to deliver exceptional performance, scalability, and resilience—powering resource management across ByteDance’s massive global infrastructure.
  • Innovate on core scheduling capabilities: Design and maintain a truly unified scheduling that powers diverse workloads (Containers & VMs, online services, offline computing, AI/ML, CPU/GPU workloads, etc) in a massive-scale resource pool.
  • Develop an intelligent scheduling system: Leverage AI models to optimize workload performance and resource utilization across heterogeneous resources—including CPU, GPU, memory, network, and power across global data centers.
  • Lead Infrastructure for Next-Gen ML Workloads: Design and drive the evolution of compute platforms purpose-built for fast, reliable, and cost-effective ML and LLM training/inference.
  • Deliver Quality and Innovation: Write high-quality, maintainable code, and stay at the forefront of open-source and research advancements in AI, ML, systems, and Serverless technologies.
Qualifications
  • Minimum Qualifications
    • B.S./M.S, degree in Computer Science, Computer Engineering or a related area with 3+ years of relevant industry experience; new graduates with Ph.D. degree and strong publication records can be an exception.
    • Solid understanding of at least one of the following fields: Unix/Linux environments, distributed and parallel systems, high-performance networking systems, developing large scale software systems.
    • Proven experience designing, architecting and building cloud and ML infrastructure related but not limited to resource management, allocation, job scheduling and monitoring.
    • Familiarity with container and orchestration technologies such as Docker and Kubernetes.
    • Proficiency in at least one major programming language such as Python, Go, C++, Rust, and Java.
  • Preferred Qualifications
    • Experience in one large scale cluster management systems, e.g., Kubernetes, Ray, Yarn, or Mesos.
    • Experience in large scale resource efficiency management and job scheduling development.
    • Project experience in application scaling, workload co-location, and isolation enhancement.
    • Experience with a public cloud provider (AWS, Azure and GCP), and their ML services (e.g., AWS SageMaker, Azure ML, GCP Vertex AI).
    • Great communication skills and the ability to work well within a team and across engineering teams.
    • Passionate about system efficiency, quality, performance and scalability.
Job Information

The base salary range for this position in the selected city is $156000 - $387600 annually.

Benefits

Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

EEO Statement (Los Angeles County)

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues; 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and 3. Exercising sound judgment.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical, dental and vision insurance
401(k) with company match
Parental leave
+6
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Tech Lead Software Engineer - AI Compute Infrastructure
Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Tech Lead, Research Scientist - DPU & AI Infra
Tech Lead, Research Scientist - DPU & AI Infra

ByteDance • San Jose (CA)

On-site
USD 244,800 - 588,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Software Engineer Graduate (Cloud Native Infrastructure)- 2026 Start (PHD)
Software Engineer Graduate (Cloud Native Infrastructure)- 2026 Start (PHD)

Pangleglobal • San Jose (CA)

On-site
USD 118,000 - 260,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
Cloud Site Reliability Engineer - DCS Cloud San Jose Regular
Cloud Site Reliability Engineer - DCS Cloud San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid personal time off
+1
Senior Software Engineer, Backend and Infrastructure
Senior Software Engineer, Backend and Infrastructure

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Senior Software Engineer, AI Infrastructure - Developer Tooling
Senior Software Engineer, AI Infrastructure - Developer Tooling

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+8
Tech Lead - Machine Learning Platform Engineer
Tech Lead - Machine Learning Platform Engineer

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+2
Senior Software Engineer- Metadata Storage
Senior Software Engineer- Metadata Storage

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+8