RL Systems Architect: Scalable AI Training & Infra

Bytedance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a Research Engineer for RL Systems & Infrastructure in San Jose. You will design end-to-end RL systems for large-scale models, build fault-tolerant infrastructure, and optimize distributed training across GPU clusters.

Collaboration with researchers to convert ideas into robust, production-ready solutions is essential. The role offers exposure to cutting-edge RL and ML workflows within a globally diverse team and a strong emphasis on reliability, performance, and

Qualifications

  • Experience with distributed systems and large-scale ML infrastructure.
  • Experience building or optimizing RL/LLM/multimodal training pipelines.
  • Solid Python/C++ programming and familiarity with PyTorch and distributed training.

Responsibilities

  • Design end-to-end RL systems for large-scale models, covering rollout, training, evaluation, and deployment.
  • Develop scalable, fault-tolerant RL infrastructure for dynamic workloads and heterogeneous compute environments.
  • Optimize distributed training across GPU clusters for throughput and stability.
  • Collaborate with researchers to translate ideas into production-grade implementations.
  • Build tooling and observability for RL training systems.

Skills

Distributed systems
Large-scale ML systems
Deep learning infrastructure
Python
C++
PyTorch
RL training workflows
GPU optimization

Tools

Distributed training frameworks

Job description

ByteDance is seeking a Research Engineer for RL Systems & Infrastructure in San Jose. You will design end-to-end RL systems for large-scale models, build fault-tolerant infrastructure, and optimize distributed training across GPU clusters.

Collaboration with researchers to convert ideas into robust, production-ready solutions is essential. The role offers exposure to cutting-edge RL and ML workflows within a globally diverse team and a strong emphasis on reliability, performance, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer – Distributed AI Infrastructure
Research Engineer – Distributed AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 254,000 - 480,000
Medical/Dental/Vision insurance
401(k) with company match
Parental leave
+5
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
Senior AI Systems Scientist - Memory & Scale
Senior AI Systems Scientist - Memory & Scale

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical Insurance
Dental Insurance
Vision Insurance
+9
Research Scientist: AI Infrastructure & Large-Scale ML
Research Scientist: AI Infrastructure & Large-Scale ML

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Infrastructure Research Engineer - Large-Scale RL Systems
Infrastructure Research Engineer - Large-Scale RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1