LLM Distributed Training Engineer (1000+ GPUs)

Hyphen Connect Limited

Oregon (WI)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hyphen Connect Limited is seeking a highly skilled LLM Pre-training & Distributed Systems Engineer to orchestrate large-scale machine learning training runs and optimize distributed infrastructure. Candidates should have a deep understanding of GPU clusters and extensive experience in systems engineering for efficient and reliable training processes.

This role involves optimizing networking and memory management during training and automating the recovery of checkpoints after failures. Interested candidates with strong expertise in parallelism, C++, CUDA, and Python are encouraged to apply.

Qualifications

  • Deep expertise in 3D parallelism: Data, Tensor, Pipeline.
  • Experience optimizing distributed systems for machine learning.
  • Proficient in systems engineering with C++, CUDA, and Python.

Responsibilities

  • Orchestrate distributed training runs across 1,000+ GPUs.
  • Optimize networking and memory management to prevent errors.
  • Automate checkpointing and failure recovery in training runs.

Skills

3D parallelism (Data, Tensor, Pipeline)
C++
CUDA
Python

Tools

PyTorch
DeepSpeed
Megatron-LM

Job description

Hyphen Connect Limited is seeking a highly skilled LLM Pre-training & Distributed Systems Engineer to orchestrate large-scale machine learning training runs and optimize distributed infrastructure. Candidates should have a deep understanding of GPU clusters and extensive experience in systems engineering for efficient and reliable training processes.

This role involves optimizing networking and memory management during training and automating the recovery of checkpoints after failures. Interested candidates with strong expertise in parallelism, C++, CUDA, and Python are encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Distributed Training Engineer (1000+ GPUs)
LLM Distributed Training Engineer (1000+ GPUs)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect • Boston (MA)

On-site
USD 120,000 - 160,000
LLM Pre-Training & Distributed Systems Architect
LLM Pre-Training & Distributed Systems Architect

Hyphen Connect • Boston (MA)

On-site
USD 120,000 - 160,000
LLM Training Architect: 1k-GPU Distributed Systems
LLM Training Architect: 1k-GPU Distributed Systems

Hyphen Connect • Seattle (WA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect • Seattle (WA)

On-site
USD 120,000 - 160,000
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
AI Infrastructure — Training Engineer (Large Model) [33251]
AI Infrastructure — Training Engineer (Large Model) [33251]

Stealth Startup • Menlo Park (CA)

On-site
USD 180,000 - 260,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000