Research Engineer – Distributed AI Infrastructure

ByteDance

San Jose (CA)

On-site

USD 254,000 - 480,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical/Dental/Vision insurance
401(k) with company match
Parental leave
Disability insurance
Life insurance
Paid holidays
Paid sick days
Paid Personal Time

Job summary

ByteDance Seed in San Jose is seeking PhD‑level researchers to advance large‑scale AI infrastructure, including distributed training, RL frameworks, and multimodal model work.

You will collaborate with researchers and engineers to translate research prototypes into production‑ready systems, optimize GPU and network throughput, and build observability tools.

This role offers bold problems, strong growth, and the chance to shape foundational AI technologies for real products.

Qualifications

  • PhD in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or a related discipline.
  • Strong background in distributed systems, large-scale machine learning systems, or deep learning infrastructure.
  • Research or hands‑on experience in training or optimizing large-scale models (e.g., LLMs, multimodal models, RL systems).
  • Understanding of parallelism strategies (e.g., data, model/tensor, pipeline, expert parallelism) and distributed training concepts.
  • Familiarity with reinforcement learning workflows such as rollout generation, policy optimization, and evaluation loops.
  • Proficiency in programming (e.g., Python and/or C++) and experience with modern ML frameworks (e.g., PyTorch and distributed training tools).

Responsibilities

  • Conduct research and development on large-scale AI infrastructure to support efficient training and post-training of foundation models, multimodal LLMs, and image/video generation models.
  • Design and optimize distributed training strategies, including data/model/tensor/pipeline/expert parallelism, computation–communication overlap, and large-scale GPU cluster scaling.
  • Prototype and improve end-to-end reinforcement learning (RL) training systems, covering rollout generation, policy optimization, evaluation, and iterative deployment workflows.
  • Build scalable and fault-tolerant infrastructure that operates reliably under dynamic workloads and heterogeneous compute environments.
  • Analyze performance bottlenecks across the training stack (e.g., networking, scheduling, GPU memory management), and develop principled optimization approaches to improve throughput, efficiency, and stability.
  • Develop tooling, monitoring, debugging, and observability frameworks to ensure reliability of large-scale training and RL systems.
  • Collaborate with researchers and engineers on system–algorithm co-design, translating research prototypes into scalable, production-ready infrastructure systems.

Skills

Distributed systems
Large-scale ML
RL workflows
Python
C++
PyTorch

Education

PhD in Computer Science, Electrical Engineering, or related field

Tools

Distributed training tools

Job description

ByteDance Seed in San Jose is seeking PhD‑level researchers to advance large‑scale AI infrastructure, including distributed training, RL frameworks, and multimodal model work.

You will collaborate with researchers and engineers to translate research prototypes into production‑ready systems, optimize GPU and network throughput, and build observability tools.

This role offers bold problems, strong growth, and the chance to shape foundational AI technologies for real products.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist: AI Infrastructure & Large-Scale ML
Research Scientist: AI Infrastructure & Large-Scale ML

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
Graduate Research Scientist, AI Infra & Compute – San Jose
Graduate Research Scientist, AI Infra & Compute – San Jose

ByteDance • San Jose (CA)

On-site
USD 180,000 - 230,000
Research Scientist, AI Infrastructure & Distributed ML
Research Scientist, AI Infrastructure & Distributed ML

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Graduate Research Scientist, AI Systems & Infrastructure
Graduate Research Scientist, AI Systems & Infrastructure

ByteDance • San Jose (CA)

On-site
USD 218,000 - 388,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+6
Research Scientist: AI Foundation Models & ML Systems
Research Scientist: AI Foundation Models & ML Systems

ByteDance • San Jose (CA)

On-site
USD 254,000 - 480,000
Graduate Research Engineer, AI Training Systems
Graduate Research Engineer, AI Training Systems

ByteDance • Seattle (WA)

On-site
USD 242,000 - 456,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
AI Infrastructure & ML Systems Intern
AI Infrastructure & ML Systems Intern

ByteDance • San Jose (CA)

On-site
USD 97,000 - 138,000
Health insurance
Housing allowance
Paid holidays
Graduate Research Scientist – AI for Infrastructure
Graduate Research Scientist – AI for Infrastructure

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Research Scientist - AI Systems & Database Innovation
Research Scientist - AI Systems & Database Innovation

ByteDance • Seattle (WA)

On-site
USD 150,000 - 190,000