Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance

Seattle (WA)

On-site

USD 232,560 - 427,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance in Seattle is seeking a Tech Lead Software Engineer for AI Compute Infrastructure to design and operate large‑scale, GPU‑accelerated inference platforms that power cutting‑edge AI workloads across the company’s data centers. You will lead architecture, tooling, and performance tuning in a fast‑paced environment.

You will collaborate with platform, ML, and product teams to deliver reliable inference systems, drive improvements in scheduling and resource management, and contribute to

Qualifications

  • BS/MS in CS/CE or related field with 5 years of relevant experience (PhD with strong systems/ML publications also considered).
  • Strong understanding of large model inference, distributed and parallel systems, and/or high‑performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++).

Responsibilities

  • Design and build large‑scale, container‑based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next‑generation cloud‑native GPU and AI accelerator infrastructure to deliver cost‑efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT‑LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high‑quality, production‑ready code that is maintainable, testable, and scalable.

Skills

Large‑scale inference
Distributed systems
Cloud infrastructure
Docker/Kubernetes
Go/Rust/Python/C++

Education

B.S./M.S. in CS/CE or related fields

Tools

Docker
Kubernetes

Job description

Tech Lead Software Engineer - AI Compute Infrastructure

Location:

Seattle

Team:

Infrastructure

Employment Type:

Regular

Job Code:

A35024

About the Team

The Inference Infrastructure team is the creator and open-source maintainer of AIBrix, a Kubernetes-native control plane for large-scale LLM inference. We are part of ByteDance’s Core Compute Infrastructure organization, responsible for designing and operating the platforms that power microservices, big data, distributed storage, machine learning training and inference, and edge computing across multi-cloud and global datacenters. With ByteDance’s rapidly growing businesses and a global fleet of machines running hundreds of millions of containers daily, we are building the next generation of cloud-native, GPU‑optimized orchestration systems. Our mission is to deliver infrastructure that is highly performant, massively scalable, cost‑efficient, and easy to use—enabling both internal and external developers to bring AI workloads from research to production at scale.

We are expanding our focus on LLM inference infrastructure to support new AI workloads, and are looking for engineers passionate about cloud‑native systems, scheduling, and GPU acceleration. You’ll work in a hyper‑scale environment, collaborate with world‑class engineers, contribute to the open‑source community, and help shape the future of AI inference infrastructure globally.

Responsibilities
  • Design and build large‑scale, container‑based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next‑generation cloud‑native GPU and AI accelerator infrastructure to deliver cost‑efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT‑LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high‑quality, production‑ready code that is maintainable, testable, and scalable.
Qualifications

Minimum Qualifications:

  • B.S./M.S. in Computer Science, Computer Engineering, or related fields with 5 years of relevant experience (Ph.D. with strong systems/ML publications also considered).
  • Strong understanding of large model inference, distributed and parallel systems, and/or high‑performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++).

Preferred Qualifications:

  • Experience contributing to or operating large‑scale cluster management systems (e.g., Kubernetes, Ray).
  • Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments.
  • Hands‑on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT‑LLM).
  • Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI).
  • Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms.
  • Excellent communication skills and ability to collaborate across global, cross‑functional teams.
  • Passion for system efficiency, performance optimization, and open‑source innovation.
Job Information

The base salary range for this position in the selected city is $232,560 - $427,500 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

Equal Employment Opportunity

For Los Angeles County (unincorporated) Candidates:

  • 1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  • 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and
  • 3. Exercising sound judgment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead Software Engineer - AI Compute Infrastructure
Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 148,000 - 301,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior Software Engineer - AI Compute Infrastructure
Senior Software Engineer - AI Compute Infrastructure

ByteDance • Seattle (WA)

On-site
USD 202,160 - 368,220
Medical, dental, vision
401(k) with company match
Paid parental leave
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 202,000 - 369,000
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD)
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD)

Pangleglobal • Seattle (WA)

On-site
USD 129,000 - 247,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental and vision insurance
401(k) with company match
Paid parental leave
+6
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
AI Infrastructure Engineer
AI Infrastructure Engineer

Netpreme • Santa Clara (CA), Boston (MA)

On-site
USD 150,000 - 210,000
Health, Dental, and Vision coverage
401(k) match
Life, Disability and AD&D insurance
+4