LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hyphen Connect Limited is seeking a highly skilled LLM Pre-training & Distributed Systems Engineer to optimize large-scale machine learning training runs. The ideal candidate will have deep expertise in GPU clusters and extensive systems engineering experience to ensure efficiency and reliability in training processes.

Key responsibilities include orchestrating distributed training runs, optimizing networking and memory management, and automating checkpointing during long training periods. If you have a passion for cutting-edge technology and system architecture, we want to hear from you.

Qualifications

  • Deep expertise in orchestrating distributed training runs across multiple GPUs.
  • Strong systems engineering skills with experience in C++, CUDA, and Python.

Responsibilities

  • Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
  • Automate checkpointing and failure recovery during month-long training runs.

Skills

3D parallelism expertise
Systems engineering (C++, CUDA, Python)

Job description

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.

Responsibilities
  • Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
  • Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
  • Automate checkpointing and failure recovery during month-long training runs.
Required Skills
  • Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
  • Strong systems engineering background (C++, CUDA, Python).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Distributed Training Engineer (1000+ GPUs)
LLM Distributed Training Engineer (1000+ GPUs)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Distributed Training Engineer (1000+ GPUs)
LLM Distributed Training Engineer (1000+ GPUs)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 290,000 - 363,000
Distributed AI Training Engineer (GPU Clusters)
Distributed AI Training Engineer (GPU Clusters)

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Training Infra Engineer for Scalable LLM Systems
Training Infra Engineer for Scalable LLM Systems

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior LLM Training Performance Engineer (Hybrid)
Senior LLM Training Performance Engineer (Hybrid)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits package
Hybrid work environment