Senior Infrastructure Engineer, AI

AI Breaking Wire

San Francisco (CA)

On-site

USD 280,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical, dental, and vision benefits
Unlimited PTO
Parental leave (flexible)
Learning stipend

Job summary

OpenAI is seeking a Senior Infrastructure Engineer to help build and scale the compute infrastructure powering state-of-the-art generative AI models in San Francisco. You will work with massive clusters of GPUs to optimize training and inference performance.

You will collaborate with research teams to remove bottlenecks, optimize networks and storage, and troubleshoot complex distributed systems across thousands of accelerators.

Qualifications

  • 5+ years of experience in software engineering with a focus on large-scale distributed systems or ML infrastructure.
  • Deep expertise in CUDA, NCCL, InfiniBand, and GPU cluster management.
  • Proficiency in Python, C++, and container orchestration tools like Kubernetes.

Responsibilities

  • Architect, build, and scale high-throughput, low-latency distributed training and inference systems.
  • Optimize cluster utilization, network topology, and storage subsystems for massive deep learning workloads.
  • Partner with research teams to remove infrastructure bottlenecks and accelerate model iteration speed.
  • Troubleshoot complex distributed systems issues across thousands of hardware accelerators.

Skills

CUDA
NCCL
InfiniBand
GPU cluster management
Python
C++
Kubernetes

Education

BS or MS in Computer Science or related technical discipline

Job description

# Senior Infrastructure Engineer, AIOpenAI## Job Description### About the RoleOpenAI is seeking a Senior Infrastructure Engineer to help build and scale the compute infrastructure powering our state-of-the-art generative AI models. You will work with massive clusters of GPUs to optimize training and inference performance.### Responsibilities- Architect, build, and scale high-throughput, low-latency distributed training and inference systems.- Optimize cluster utilization, network topology, and storage subsystems for massive deep learning workloads.- Partner with research teams to remove infrastructure bottlenecks and accelerate model iteration speed.- Troubleshoot complex distributed systems issues across thousands of hardware accelerators.### Requirements- 5+ years of experience in software engineering with a focus on large-scale distributed systems or ML infrastructure.- Deep expertise in CUDA, NCCL, InfiniBand, and GPU cluster management.- Proficiency in Python, C++, and container orchestration tools like Kubernetes.- BS or MS in Computer Science or a related technical discipline.### Benefits- Industry-leading compensation including equity.- Top-tier medical, dental, and vision benefits.- Unlimited paid time off and flexible parental leave.- Annual learning and development stipend.## Skills & Tagspythonc++cudakubernetesdistributed-systemsinfrastructure## Job DetailsFull-timeSan Francisco, CA$280k – $400k USDPosted July 30, 2026Expires September 28, 2026
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, Infrastructure
Senior Machine Learning Engineer, Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior Software Engineer, AI Infrastructure $126,000 - $189,000 Posted 3 hours ago
Senior Software Engineer, AI Infrastructure $126,000 - $189,000 Posted 3 hours ago

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Applied Machine Learning Engineer
Applied Machine Learning Engineer

AI Breaking Wire • San Francisco (CA)

Hybrid
USD 250,000 - 380,000
Equity
Medical, dental, and vision
Unlimited PTO
+2
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

CV in • Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits