Lead, AI Compute Infra for Scalable LLM Inference

ByteDance

Seattle (WA)

On-site

USD 232,560 - 427,500

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ByteDance in Seattle is seeking a Tech Lead Software Engineer for AI Compute Infrastructure to design and operate large‑scale, GPU‑accelerated inference platforms that power cutting‑edge AI workloads across the company’s data centers. You will lead architecture, tooling, and performance tuning in a fast‑paced environment.

You will collaborate with platform, ML, and product teams to deliver reliable inference systems, drive improvements in scheduling and resource management, and contribute to

Qualifications

  • BS/MS in CS/CE or related field with 5 years of relevant experience (PhD with strong systems/ML publications also considered).
  • Strong understanding of large model inference, distributed and parallel systems, and/or high‑performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++).

Responsibilities

  • Design and build large‑scale, container‑based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next‑generation cloud‑native GPU and AI accelerator infrastructure to deliver cost‑efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT‑LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high‑quality, production‑ready code that is maintainable, testable, and scalable.

Skills

Large‑scale inference
Distributed systems
Cloud infrastructure
Docker/Kubernetes
Go/Rust/Python/C++

Education

B.S./M.S. in CS/CE or related fields

Tools

Docker
Kubernetes

Job description

ByteDance in Seattle is seeking a Tech Lead Software Engineer for AI Compute Infrastructure to design and operate large‑scale, GPU‑accelerated inference platforms that power cutting‑edge AI workloads across the company’s data centers. You will lead architecture, tooling, and performance tuning in a fast‑paced environment.

You will collaborate with platform, ML, and product teams to deliver reliable inference systems, drive improvements in scheduling and resource management, and contribute to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Tech Lead - AI Infra & DPU Research Scientist
Tech Lead - AI Infra & DPU Research Scientist

ByteDance • Seattle (WA)

On-site
USD 256,500 - 427,500
Medical insurance
Dental insurance
Vision insurance
+8
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
AI Infrastructure Lead & Research Architect
AI Infrastructure Lead & Research Architect

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical insurance
401(k) match
Parental leave
+3
AI Compute & DPU Research Scientist
AI Compute & DPU Research Scientist

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental and vision insurance
401(k) with company match
Paid parental leave
+6
Lead, AI Inference & GPU Strategy
Lead, AI Inference & GPU Strategy

DigitalOcean • Seattle (WA)

Hybrid
USD 218,000 - 273,000
Lead AI Infra Engineer: GPU Compute & ML Pipelines, Equity
Lead AI Infra Engineer: GPU Compute & ML Pipelines, Equity

Fuel Talent • Seattle (WA)

On-site
USD 180,000 - 210,000
Tech Lead: AI Infra & DPU Research
Tech Lead: AI Infra & DPU Research

ByteDance • San Jose (CA)

On-site
USD 244,800 - 588,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior DPU & AI Infrastructure Scientist
Senior DPU & AI Infrastructure Scientist

ByteDance • Seattle (WA)

On-site
USD 202,160 - 368,220
Medical insurance
Dental insurance
Vision insurance
+8