Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)

ByteDance

Seattle (WA)

On-site

USD 148,200 - 300,960

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
Paid holidays
Paid sick days

Job summary

ByteDance in Seattle is seeking engineers to join the Inference Infrastructure team, focusing on large-scale AI workloads. As a part of a hyper-scale environment, you'll design orchestration systems, ensuring performance and resilience. Candidates should have a PhD in a relevant field with experience in cloud infrastructure and container technologies. Offering competitive salaries from $148,200 to $300,960 with comprehensive benefits including medical insurance and a retirement plan, this role provides a platform for innovation and growth.

Qualifications

  • Completion of a PhD in Software Development, Computer Science, or related field.
  • Strong understanding of distributed and parallel systems.
  • Experience with cloud or ML infrastructure.

Responsibilities

  • Design and build large-scale orchestration systems.
  • Collaborate to deliver inference solutions using various LLM engines.
  • Write maintainable, production-ready code.

Skills

PhD in Software Development
Understanding of large model inference
Experience building cloud/ML infrastructure
Knowledge of container technologies
Proficiency in programming languages

Education

PhD degree in related discipline

Tools

Docker
Kubernetes

Job description

About the Team

The Inference Infrastructure team is the creator and open‑source maintainer of AIBrix, a Kubernetes‑native control plane for large‑scale LLM inference. We are part of ByteDance’s Core Compute Infrastructure organization, responsible for designing and operating the platforms that power microservices, big data, distributed storage, machine learning training and inference, and edge computing across multi‑cloud and global datacenters.

With ByteDance’s rapidly growing businesses and a global fleet of machines running hundreds of millions of containers daily, we are building the next generation of cloud‑native, GPU‑optimized orchestration systems. Our mission is to deliver infrastructure that is highly performant, massively scalable, cost‑efficient, and easy to use—enabling both internal and external developers to bring AI workloads from research to production at scale.

We are expanding our focus on LLM inference infrastructure to support new AI workloads, and are looking for engineers passionate about cloud‑native systems, scheduling, and GPU acceleration. You’ll work in a hyper‑scale environment, collaborate with world‑class engineers, contribute to the open‑source community, and help shape the future of AI inference infrastructure globally.

We are looking for talented individuals to join our team in 2026. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.

Successful candidates must be able to commit to an onboarding date by the end of year 2026. Please state your availability and graduation date clearly in your resume.

Responsibilities
  • Design and build large‑scale, container‑based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next‑generation cloud‑native GPU and AI accelerator infrastructure to deliver cost‑efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT‑LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high‑quality, production‑ready code that is maintainable, testable, and scalable.
Minimum Qualifications
  • Individuals who are completing or have recently completed a PhD degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • Strong understanding of large model inference, distributed and parallel systems, and/or high‑performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++).
Preferred Qualifications
  • Experience contributing to or operating large‑scale cluster management systems (e.g., Kubernetes, Ray).
  • Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments.
  • Hands‑on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT‑LLM).
  • Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI).
  • Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms.
  • Excellent communication skills and ability to collaborate across global, cross‑functional teams.
  • Passion for system efficiency, performance optimization, and open‑source innovation.
Compensation and Benefits

The base salary range for this position in the selected city is $148,200 - $300,960 annually. Compensation may vary outside of this range depending on a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day‑one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year, and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

Equality and Fair Employment
  • Qualifying applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.
  • Job duties that may be impacted include:
    • Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
    • Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
    • Exercising sound judgment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead Software Engineer - AI Compute Infrastructure
Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 202,000 - 369,000
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD)
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PHD)

Pangleglobal • Seattle (WA)

On-site
USD 129,000 - 247,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Research Scientist - AI Compute & DPU - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental and vision insurance
401(k) with company match
Paid parental leave
+6
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Software Engineer Graduate (AI Infra Compute) - 2027 Start
Software Engineer Graduate (AI Infra Compute) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Software Engineer Graduate (AI Infra Compute) - 2027 Start
Software Engineer Graduate (AI Infra Compute) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 120,000 - 180,000
Research Scientist Graduate (Infrastructure System Lab)- 2026 Start (PHD)
Research Scientist Graduate (Infrastructure System Lab)- 2026 Start (PHD)

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
+2
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

Bytedance • San Jose (CA)

On-site
USD 128,000 - 256,000