PhD Research Intern — Cloud AI Systems & HPC

ByteDance

San Jose (CA)

On-site

USD 68,000 - 97,000

Full time

30 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Life insurance
Wellbeing benefits
Paid holidays
Housing allowance

Job summary

ByteDance invites PhD students to join the DPU team for a 12-week Summer 2027 internship in San Jose, CA. You will design and build large-scale, container-based cluster management and GPU/AI accelerator infrastructure, collaborating across teams to deliver inference solutions.

We value hands-on contributors with strong background in distributed systems, cloud ML infrastructure, and experience with Docker, Kubernetes, and modern ML frameworks; excellent communication across global teams is a plus.

Qualifications

  • Currently pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • Able to commit to working for 12 weeks during Summer 2027
  • Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems.
  • Hands‑on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration.
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes).
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++).

Responsibilities

  • Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience.
  • Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms.
  • Collaborate across teams to deliver world‑class inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines.
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems.
  • Write high-quality, production-ready code that is maintainable, testable, and scalable.

Skills

Go
Rust
Python
C++
Distributed systems
Large model inference

Education

PhD candidate (CS/CE/EE)

Tools

Docker
Kubernetes
TensorRT
Ray
SGLang
vLLM

Job description

ByteDance invites PhD students to join the DPU team for a 12-week Summer 2027 internship in San Jose, CA. You will design and build large-scale, container-based cluster management and GPU/AI accelerator infrastructure, collaborating across teams to deliver inference solutions.

We value hands-on contributors with strong background in distributed systems, cloud ML infrastructure, and experience with Docker, Kubernetes, and modern ML frameworks; excellent communication across global teams is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PhD Cloud & AI Systems Research Intern
PhD Cloud & AI Systems Research Intern

ByteDance • San Jose (CA)

On-site
USD 68,000 - 97,000
Health insurance
Wellbeing benefits
Paid holidays
+2
Research Intern, Cloud & AI Systems
Research Intern, Cloud & AI Systems

ByteDance • Seattle (WA)

On-site
USD 55,000 - 83,000
AI Compute Research Intern: Cloud & GPUs (PhD, Summer 2027)
AI Compute Research Intern: Cloud & GPUs (PhD, Summer 2027)

ByteDance • Seattle (WA)

On-site
USD 65,000 - 92,000
Health insurance
Housing allowance
Paid holidays
PhD Research Intern — AI Infra & Cloud Compute
PhD Research Intern — AI Infra & Cloud Compute

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 34,000 - 48,000
DPU-Driven Cloud Acceleration Research Intern
DPU-Driven Cloud Acceleration Research Intern

ByteDance • Seattle (WA)

On-site
USD 201,000 - 268,000
AI Systems & Cloud Infrastructure Intern
AI Systems & Cloud Infrastructure Intern

ByteDance • San Jose (CA)

On-site
USD 50,000 - 83,000
Health insurance
Life insurance
Wellbeing benefits
+3
Cloud Acceleration Research Intern (DPU & AI Infra) - 2027 Start (PhD)
Cloud Acceleration Research Intern (DPU & AI Infra) - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 201,000 - 268,000
PhD AI Infra Research Intern — Compute Platform
PhD AI Infra Research Intern — Compute Platform

ByteDance • San Jose (CA)

On-site
USD 68,000 - 97,000
Health insurance from day one
Wellbeing benefits
10 paid holidays per year
+1
PhD AI Systems Intern - Large-Scale ML Infrastructure
PhD AI Systems Intern - Large-Scale ML Infrastructure

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 34,000 - 55,000
Graduate Research Scientist, DPU & AI/ML Infrastructure
Graduate Research Scientist, DPU & AI/ML Infrastructure

ByteDance • Seattle (WA)

On-site
USD 180,000 - 240,000