LLM Training Architect: 1k-GPU Distributed Systems
Hyphen Connect
Seattle (WA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Hyphen Connect is looking for an experienced LLM Pre-training & Distributed Systems Engineer in Seattle, USA. This role involves orchestrating large-scale machine learning training runs across 1,000+ GPUs, optimizing distributed infrastructure, and ensuring efficient training processes. The ideal candidate should possess deep expertise in 3D parallelism and a strong systems engineering background with C++, CUDA, and Python. Join us to push the boundaries of machine learning infrastructure.
Qualifications
Deep expertise in 3D parallelism required.
Strong background in systems engineering with C++, CUDA, and Python.
Responsibilities
Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
Optimize networking (InfiniBand/RDMA) and memory management.
Automate checkpointing and failure recovery for long training runs.
Skills
3D parallelism (Data, Tensor, Pipeline)
Systems engineering (C++, CUDA, Python)
Job description
Hyphen Connect is looking for an experienced LLM Pre-training & Distributed Systems Engineer in Seattle, USA. This role involves orchestrating large-scale machine learning training runs across 1,000+ GPUs, optimizing distributed infrastructure, and ensuring efficient training processes. The ideal candidate should possess deep expertise in 3D parallelism and a strong systems engineering background with C++, CUDA, and Python. Join us to push the boundaries of machine learning infrastructure.