Senior Machine Learning Engineer, Infrastructure

AI Breaking Wire

San Francisco, Northern (CA, KY)

On-site

USD 320,000 - 500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical/Dental/Vision coverage
Unlimited PTO
Wellness stipends

Job summary

OpenAI is seeking a Senior ML Infrastructure Engineer to architect and scale large-scale training and inference systems. You will work with researchers to optimize performance, throughput, and reliability across tens of thousands of GPUs.

The role requires deep expertise in distributed systems, CUDA, Python, C++, and modern orchestration tools, with collaboration across hardware vendors for next-gen accelerators.

Qualifications

  • BS/MS/PhD in CS or related field.
  • 5+ years of software engineering experience with distributed systems.
  • Proficiency in C++, Python, CUDA, and modern cluster orchestration tools (Kubernetes, Ray).
  • Deep understanding of networking (InfiniBand, RoCE) and distributed training frameworks Megatron-LM, DeepSpeed.

Responsibilities

  • Architect, build, and scale distributed training and inference infrastructure.
  • Diagnose and resolve complex performance bottlenecks in multi-node GPU clusters.
  • Develop tools and automation for CI/CD, deployment, and monitoring of ML workloads.
  • Collaborate with hardware vendors to prototype and integrate next-gen accelerator chips.

Skills

Python
C++
CUDA
Kubernetes
Distributed systems
Infrastructure

Education

B.S./M.S./Ph.D. in Computer Science or related field

Tools

Megatron-LM
DeepSpeed
Ray

Job description

# Senior Machine Learning Engineer, InfrastructureOpenAI## Job Description### About the RoleWe are seeking a Senior ML Infrastructure Engineer to build the backbone of our large-scale training and inference systems. You will work closely with researchers to optimize the performance, throughput, and reliability of cluster operations running across tens of thousands of GPUs.### Responsibilities- Architect, build, and scale distributed training and inference infrastructure.- Diagnose and resolve complex performance bottlenecks in multi-node GPU clusters.- Develop tools and automation for continuous integration, deployment, and monitoring of ML workloads.- Collaborate with hardware vendors to prototype and integrate next-gen accelerator chips.### Requirements- B.S., M.S., or Ph.D. in Computer Science or related field.- 5+ years of software engineering experience with deep expertise in distributed systems.- Proficiency in C++, Python, CUDA, and modern cluster orchestration tools (Kubernetes, Ray).- Deep understanding of networking protocols (InfiniBand, RoCE) and distributed training frameworks (Megatron-LM, DeepSpeed).### Benefits- Industry-leading compensation and equity.- Full medical, dental, and vision coverage with 100% premium coverage options.- Unlimited paid time off and flexible remote work policies.- Wellness stipends and learning budgets.## Skills & Tagspythonc++cudakubernetesdistributed-systemsinfrastructure## Job DetailsFull-timeSan Francisco, CA$320k – $500k USDPosted August 13, 2026Expires October 12, 2026
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer, AI
Senior Infrastructure Engineer, AI

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
AI Infrastructure Engineer
AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Software Engineer, Machine Learning Infrastructure
Software Engineer, Machine Learning Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Applied Machine Learning Engineer
Applied Machine Learning Engineer

AI Breaking Wire • San Francisco (CA)

Hybrid
USD 250,000 - 380,000
Equity
Medical, dental, and vision
Unlimited PTO
+2
Senior Software Engineer, AI Infrastructure $126,000 - $189,000 Posted 3 hours ago
Senior Software Engineer, AI Infrastructure $126,000 - $189,000 Posted 3 hours ago

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior Machine Learning Engineer, Core Systems
Senior Machine Learning Engineer, Core Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Data Infrastructure Engineer — GPU-Scale Datasets & APIs
Data Infrastructure Engineer — GPU-Scale Datasets & APIs

Slope • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 400,000