ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 170,000 - 250,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
Catered lunches & on-site wellness

Job summary

OpenAI in San Francisco is seeking an experienced engineer to architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs. You will optimize GPU cluster utilization and memory management for multi-node training, partnering with research teams to accelerate experiments.

The role requires deep expertise in distributed systems, networking, and high-throughput computing, plus strong coding in C++, Python, and CUDA, with hands-on experience on large

Qualifications

  • 3+ years of industry experience in distributed systems or related field.
  • Deep expertise in distributed systems, networking, and high-throughput computing.
  • Strong programming in C++, Python, and CUDA.

Responsibilities

  • Architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs.
  • Optimize GPU cluster utilization, network topologies, and memory management for multi-node training jobs.
  • Collaborate with research teams to streamline experiment velocity and scale up model architectures.
  • Troubleshoot complex distributed system bottlenecks, hardware failures, and performance regressions.

Skills

Distributed systems
Networking
High-throughput computing
C++
Python
CUDA
Kubernetes
Slurm

Education

B.S./M.S./Ph.D. in CS or related

Tools

Kubernetes
Slurm

Job description

OpenAI in San Francisco is seeking an experienced engineer to architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs. You will optimize GPU cluster utilization and memory management for multi-node training, partnering with research teams to accelerate experiments.

The role requires deep expertise in distributed systems, networking, and high-throughput computing, plus strong coding in C++, Python, and CUDA, with hands-on experience on large

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Software Engineer, Machine Learning Infrastructure
Software Engineer, Machine Learning Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000