Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 320,000 - 500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical/Dental/Vision coverage
Unlimited PTO
Wellness stipends

Job summary

OpenAI is seeking a Senior ML Infrastructure Engineer to architect and scale large-scale training and inference systems. You will work with researchers to optimize performance, throughput, and reliability across tens of thousands of GPUs.

The role requires deep expertise in distributed systems, CUDA, Python, C++, and modern orchestration tools, with collaboration across hardware vendors for next-gen accelerators.

Qualifications

  • BS/MS/PhD in CS or related field.
  • 5+ years of software engineering experience with distributed systems.
  • Proficiency in C++, Python, CUDA, and modern cluster orchestration tools (Kubernetes, Ray).
  • Deep understanding of networking (InfiniBand, RoCE) and distributed training frameworks Megatron-LM, DeepSpeed.

Responsibilities

  • Architect, build, and scale distributed training and inference infrastructure.
  • Diagnose and resolve complex performance bottlenecks in multi-node GPU clusters.
  • Develop tools and automation for CI/CD, deployment, and monitoring of ML workloads.
  • Collaborate with hardware vendors to prototype and integrate next-gen accelerator chips.

Skills

Python
C++
CUDA
Kubernetes
Distributed systems
Infrastructure

Education

B.S./M.S./Ph.D. in Computer Science or related field

Tools

Megatron-LM
DeepSpeed
Ray

Job description

OpenAI is seeking a Senior ML Infrastructure Engineer to architect and scale large-scale training and inference systems. You will work with researchers to optimize performance, throughput, and reliability across tens of thousands of GPUs.

The role requires deep expertise in distributed systems, CUDA, Python, C++, and modern orchestration tools, with collaboration across hardware vendors for next-gen accelerators.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Senior Machine Learning Engineer, Infrastructure
Senior Machine Learning Engineer, Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Remote ML Infrastructure Engineer: GPU Scale & Platform
Remote ML Infrastructure Engineer: GPU Scale & Platform

Bright Vision Technologies • Sammamish (WA)

On-site
USD 100,000 - 150,000
Senior DL Infra Engineer: Scalable GPU AI Training
Senior DL Infra Engineer: Scalable GPU AI Training

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000