Research Engineer - LLM Infra training - Seed Infra San Jose Regular

ByteDance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
Short-term and long-term disability coverage
Life insurance
10 paid holidays
10 paid sick days
17 days of Paid Personal Time

Job summary

ByteDance is hiring for a position focused on large-scale LLM training infrastructure in San Jose, California. The successful candidate will conduct research and development on distributed training strategies, optimize performance across hardware, and analyze existing systems for enhancements. The role demands extensive experience with ML systems, programming in Python or C++, and understanding of parallelism strategies. A competitive salary ranging from $244,800 to $450,000 annually, plus benefits, is offered. Opportunities for bonuses and incentives are also available.

Qualifications

  • Experience with large-scale distributed training for LLMs.
  • Strong programming skills in Python and/or C++.
  • Strong background in ML systems / training infrastructure development.
  • Proficiency in parallelism strategies (DDP, FSDP).
  • Solid understanding of training stack internals (PyTorch, CUDA, NCCL).
  • Experience in performance optimization (memory, communication, throughput).

Responsibilities

  • Conduct research and development on large-scale LLM training infrastructure.
  • Design and optimize distributed training strategies for LLMs.
  • Investigate system reliability and resilience techniques.
  • Research and optimize network, scheduling, and GPU memory management.
  • Analyze performance bottlenecks in exascale training systems.
  • Translate research ideas into scalable AI infrastructure solutions.

Skills

Large-scale distributed training for LLMs
Programming skills in Python and/or C++
ML systems / training infrastructure development
Proficiency in parallelism strategies (DDP, FSDP)
Understanding of training stack internals (PyTorch, CUDA, NCCL)
Performance optimization (memory, communication, throughput)

Job description

Location: San Jose

Team: Technology

Employment Type: Regular

Job Code: A72032

Responsibilities

Team Information: The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

  • Conduct research and development on large-scale LLM training infrastructure and efficiency.
  • Design and optimize distributed training strategies for LLMs, including parallelism schemes, computation and communication optimization, and throughput scaling on large GPU clusters.
  • Investigate system reliability and resilience techniques, such as fast checkpointing, fault tolerance, and failure diagnosis for long-running training workloads.
  • Research and optimize network, scheduling, and GPU memory management across the training stack, driving cross-layer performance improvements.
  • Analyze performance bottlenecks in exascale training systems and propose principled, data-driven optimization methods.
  • Bridge cutting-edge research and large-scale production deployment by translating research ideas into scalable, real-world AI infrastructure solutions.
Qualifications

Minimum Qualifications:

  • Experience with large-scale distributed training for LLMs.
  • Strong programming skills in Python and/or C++.
  • Strong background in ML systems / training infrastructure development.
  • Proficiency in parallelism strategies (DDP, FSDP, model/pipeline/expert parallelism).
  • Solid understanding of training stack internals (PyTorch, CUDA, NCCL).
  • Experience in performance optimization (memory, communication, throughput).

Preferred Qualifications:

  • Hands‑on experience with distributed training frameworks and large-scale LLM infrastructure.
  • Experience leading or mentoring engineering teams or cross‑functional projects.
  • Publications in top‑tier AI, systems, or HPC conferences (ICML, OSDI, SOSP, NSDI, SIGCOMM, MLSys) or strong open‑source contributions.
  • Familiarity with benchmarking AI accelerators or large‑scale LLM evaluation (e.g., ByteMLPerf).
Job Information

The base salary range for this position in the selected city is $244,800 - $450,000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the total package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The company reserves the right to modify or change these benefits programs at any time, with or without notice.

Legal and EEO Statements

For Los Angeles County (unincorporated) candidates:

  1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
  3. Exercising sound judgment.

We uphold the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws.

Reasonable Accommodation: ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - LLM Training Infrastructure - Seed Infra Seattle Regular
Research Engineer - LLM Training Infrastructure - Seed Infra Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 232,000 - 428,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior Research Scientist - Machine Learning System
Senior Research Scientist - Machine Learning System

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular
Technology - Infrastructure Global Frontier Tech Recruitment Program - 2027 Grad San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 212,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior Research Scientist/Engineer - AI Infrastructure
Senior Research Scientist/Engineer - AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical Insurance
Dental Insurance
Vision Insurance
+9
Applied Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program[...]
Applied Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program[...]

ByteDance • San Jose (CA)

On-site
USD 212,800 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Seattle Regular
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 232,000 - 428,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Research Engineer / Scientist - Storage for LLM Technology - Infrastructure San Jose Regular
Research Engineer / Scientist - Storage for LLM Technology - Infrastructure San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 156,000 - 388,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Research Scientist Graduate (Seed-LLM) - 2027 Start (PhD)
Research Scientist Graduate (Seed-LLM) - 2027 Start (PhD)

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+2
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) San Jose Regular
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) San Jose Regular

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+1
Tech Lead, Research Scientist/Engineer - AI Infrastructure
Tech Lead, Research Scientist/Engineer - AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical insurance
401(k) match
Parental leave
+3