Software Engineer, Machine Learning Infrastructure

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 170,000 - 250,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
Catered lunches & on-site wellness

Job summary

OpenAI in San Francisco is seeking an experienced engineer to architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs. You will optimize GPU cluster utilization and memory management for multi-node training, partnering with research teams to accelerate experiments.

The role requires deep expertise in distributed systems, networking, and high-throughput computing, plus strong coding in C++, Python, and CUDA, with hands-on experience on large

Qualifications

  • 3+ years of industry experience in distributed systems or related field.
  • Deep expertise in distributed systems, networking, and high-throughput computing.
  • Strong programming in C++, Python, and CUDA.

Responsibilities

  • Architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs.
  • Optimize GPU cluster utilization, network topologies, and memory management for multi-node training jobs.
  • Collaborate with research teams to streamline experiment velocity and scale up model architectures.
  • Troubleshoot complex distributed system bottlenecks, hardware failures, and performance regressions.

Skills

Distributed systems
Networking
High-throughput computing
C++
Python
CUDA
Kubernetes
Slurm

Education

B.S./M.S./Ph.D. in CS or related

Tools

Kubernetes
Slurm

Job description

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity. Our infrastructure team builds the massive distributed systems required to train and serve frontier models.


Responsibilities


  • Architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs.

  • Optimize GPU cluster utilization, network topologies, and memory management for multi-node training jobs.

  • Collaborate with research teams to streamline experiment velocity and scale up next-generation model architectures.

  • Troubleshoot complex distributed system bottlenecks, hardware failures, and performance regressions in production.


Requirements


  • B.S., M.S., or Ph.D. in Computer Science or related technical field with 3+ years of industry experience.

  • Deep expertise in distributed systems, networking (InfiniBand/RoCE), and high-throughput computing.

  • Strong programming skills in C++, Python, and CUDA.

  • Experience working with large-scale GPU clusters (NVIDIA A100/H100/B200) and orchestration tools (Kubernetes, Slurm).


Benefits


  • Industry-leading compensation and equity packages.

  • Unlimited PTO and flexible work arrangements.

  • Top-tier medical, dental, and vision benefits.

  • Catered daily lunches and on-site wellness programs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, Infrastructure
Senior Machine Learning Engineer, Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior Infrastructure Engineer, AI
Senior Infrastructure Engineer, AI

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Software Engineer, Platform Systems
Software Engineer, Platform Systems

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Systems Generalist, GPT Infrastructure
Systems Generalist, GPT Infrastructure

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 445,000
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
AI Infrastructure Engineer
AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Data Infrastructure Engineer — GPU-Scale Datasets & APIs
Data Infrastructure Engineer — GPU-Scale Datasets & APIs

Slope • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Training Performance Engineer
Training Performance Engineer

Slope • San Francisco (CA)

On-site
USD 250,000 - 460,000
Relocation assistance
Flexible working hours
Collaborative work environment
Applied Machine Learning Engineer
Applied Machine Learning Engineer

AI Breaking Wire • San Francisco (CA)

Hybrid
USD 250,000 - 380,000
Equity
Medical, dental, and vision
Unlimited PTO
+2