Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Day‑one health benefits
401(k) with company match
Parental leave
Disability coverage

Job summary

ByteDance is seeking a Tech Lead Software Engineer to shape AI compute infrastructure in San Jose, driving large‑scale inference platforms and cloud‑native orchestration across multi‑cloud datacenters.

You’ll work on Kubernetes‑native control planes for LLM workloads, collaborate with global teams, and contribute to production‑grade systems using cutting‑edge ML tooling and open‑source technologies.

Qualifications

  • BS/MS in Computer Science, Computer Engineering, or related fields with 5 years of relevant experience (PhD with strong systems/ML publications also considered).
  • Strong understanding of large model inference, distributed and parallel systems, and/or high‑performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • Proficiency in Go, Rust, Python, or C++.

Responsibilities

  • Design and build large‑scale, container‑based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next‑generation cloud‑native GPU and AI accelerator infrastructure to deliver cost‑efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT‑LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high‑quality, production‑ready code that is maintainable, testable, and scalable.

Skills

Large model inference
Distributed systems
High-performance networking
Cloud/ML infrastructure
Docker
Kubernetes
Go
Rust
Python
C++

Education

B.S./M.S. in CS/CE or related fields

Tools

Docker
Kubernetes
CUDA
TensorRT-LLM
vLLM

Job description

Tech Lead Software Engineer - AI Compute Infrastructure

Location: San Jose

Team: Infrastructure

Employment Type: Regular

Job Code: A201019

Responsibilities

About the Team: The Inference Infrastructure team is the creator and open‑source maintainer of AIBrix, a Kubernetes‑native control plane for large‑scale LLM inference. We are part of ByteDance’s Core Compute Infrastructure organization, responsible for designing and operating the platforms that power microservices, big data, distributed storage, machine learning training and inference, and edge computing across multi‑cloud and global datacenters. With ByteDance’s rapidly growing businesses and a global fleet of machines running hundreds of millions of containers daily, we are building the next generation of cloud‑native, GPU‑optimized orchestration systems. Our mission is to deliver infrastructure that is highly performant, massively scalable, cost‑efficient, and easy to use—enabling both internal and external developers to bring AI workloads from research to production at scale. We are expanding our focus on LLM inference infrastructure to support new AI workloads, and are looking for engineers passionate about cloud‑native systems, scheduling, and GPU acceleration. You’ll work in a hyper‑scale environment, collaborate with world‑class engineers, contribute to the open‑source community, and help shape the future of AI inference infrastructure globally.

  • Design and build large‑scale, container‑based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next‑generation cloud‑native GPU and AI accelerator infrastructure to deliver cost‑efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT‑LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high‑quality, production‑ready code that is maintainable, testable, and scalable.
Qualifications
Minimum Qualifications
  • B.S./M.S. in Computer Science, Computer Engineering, or related fields with 5 years of relevant experience (Ph.D. with strong systems/ML publications also considered).
  • Strong understanding of large model inference, distributed and parallel systems, and/or high‑performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++).
Preferred Qualifications
  • Experience contributing to or operating large‑scale cluster management systems (e.g., Kubernetes, Ray).
  • Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments.
  • Hands‑on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT‑LLM).
  • Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI).
  • Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms.
  • Excellent communication skills and ability to collaborate across global, cross‑functional teams.
  • Passion for system efficiency, performance optimization, and open‑source innovation.
Job Information

The base salary range for this position in the selected city is $244,800 - $450,000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives and restricted stock units.

Benefits

Employees have day‑one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of paid personal time (prorated upon hire with increasing accruals by tenure). The company reserves the right to modify or change these benefits programs at any time, with or without notice.

Equal Employment Opportunity and Reasonable Accommodation

Qualified applicants with arrest or conviction records will be considered in accordance with all federal, state, and local laws. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties:

  1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
  3. Exercising sound judgment.

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 148,000 - 301,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 202,000 - 369,000
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental and vision insurance
401(k) with company match
Paid parental leave
+6
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD)
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD)

Pangleglobal • Seattle (WA)

On-site
USD 129,000 - 247,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Software Engineer Graduate (AI Infra Compute) - 2027 Start
Software Engineer Graduate (AI Infra Compute) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Software Engineer Graduate (Cloud Native Infrastructure)- 2026 Start (PHD)
Software Engineer Graduate (Cloud Native Infrastructure)- 2026 Start (PHD)

Pangleglobal • San Jose (CA)

On-site
USD 118,000 - 260,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
Tech Lead, Research Scientist - DPU & AI Infra
Tech Lead, Research Scientist - DPU & AI Infra

ByteDance • San Jose (CA)

On-site
USD 244,800 - 588,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)
Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Medical insurance
401(k) plan with company match
Parental leave
+6
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior Research Scientist - Machine Learning System
Senior Research Scientist - Machine Learning System

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5