LLM Training Infra Engineer - Distributed Systems

ByteDance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
Short-term and long-term disability coverage
Life insurance
10 paid holidays
10 paid sick days
17 days of Paid Personal Time

Job summary

ByteDance is hiring for a position focused on large-scale LLM training infrastructure in San Jose, California. The successful candidate will conduct research and development on distributed training strategies, optimize performance across hardware, and analyze existing systems for enhancements. The role demands extensive experience with ML systems, programming in Python or C++, and understanding of parallelism strategies. A competitive salary ranging from $244,800 to $450,000 annually, plus benefits, is offered. Opportunities for bonuses and incentives are also available.

Qualifications

  • Experience with large-scale distributed training for LLMs.
  • Strong programming skills in Python and/or C++.
  • Strong background in ML systems / training infrastructure development.
  • Proficiency in parallelism strategies (DDP, FSDP).
  • Solid understanding of training stack internals (PyTorch, CUDA, NCCL).
  • Experience in performance optimization (memory, communication, throughput).

Responsibilities

  • Conduct research and development on large-scale LLM training infrastructure.
  • Design and optimize distributed training strategies for LLMs.
  • Investigate system reliability and resilience techniques.
  • Research and optimize network, scheduling, and GPU memory management.
  • Analyze performance bottlenecks in exascale training systems.
  • Translate research ideas into scalable AI infrastructure solutions.

Skills

Large-scale distributed training for LLMs
Programming skills in Python and/or C++
ML systems / training infrastructure development
Proficiency in parallelism strategies (DDP, FSDP)
Understanding of training stack internals (PyTorch, CUDA, NCCL)
Performance optimization (memory, communication, throughput)

Job description

ByteDance is hiring for a position focused on large-scale LLM training infrastructure in San Jose, California. The successful candidate will conduct research and development on distributed training strategies, optimize performance across hardware, and analyze existing systems for enhancements. The role demands extensive experience with ML systems, programming in Python or C++, and understanding of parallelism strategies. A competitive salary ranging from $244,800 to $450,000 annually, plus benefits, is offered. Opportunities for bonuses and incentives are also available.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Training Infrastructure Research Engineer
LLM Training Infrastructure Research Engineer

ByteDance • Seattle (WA)

On-site
USD 232,000 - 428,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Graduate Infrastructure Systems Research Scientist (AI/LLM)
Graduate Infrastructure Systems Research Scientist (AI/LLM)

ByteDance • San Jose (CA)

On-site
USD 156,000 - 388,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
+2
LLM Storage Systems Research Engineer
LLM Storage Systems Research Engineer

ByteDance • San Jose (CA)

On-site
USD 156,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
ML Systems Engineer: Distributed LLM Training & Inference
ML Systems Engineer: Distributed LLM Training & Inference

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 200,000 - 251,000
Comprehensive health coverage
Equity-based compensation
Retirement benefits
+3
LLM Training & Inference Scientist (GPU-Optimized)
LLM Training & Inference Scientist (GPU-Optimized)

ByteDance • San Jose (CA)

On-site
USD 212,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
AI Infrastructure & LLM Systems Research Scientist
AI Infrastructure & LLM Systems Research Scientist

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
AI Infra Engineer: LLM Training & Inference
AI Infra Engineer: LLM Training & Inference

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Infra & LLM Systems Engineer — Founding Team
Infra & LLM Systems Engineer — Founding Team

Amadeus Search • San Francisco (CA)

Hybrid
USD 170,000 - 220,000
Research Engineer - LLM Infra training - Seed Infra San Jose Regular
Research Engineer - LLM Infra training - Seed Infra San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5
Graduate Research Scientist, LLM & AI Infrastructure
Graduate Research Scientist, LLM & AI Infrastructure

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000