Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire

San Francisco (CA)

On-site

USD 280,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical, dental, and vision benefits
Unlimited PTO
Parental leave (flexible)
Learning stipend

Job summary

OpenAI is seeking a Senior Infrastructure Engineer to help build and scale the compute infrastructure powering state-of-the-art generative AI models in San Francisco. You will work with massive clusters of GPUs to optimize training and inference performance.

You will collaborate with research teams to remove bottlenecks, optimize networks and storage, and troubleshoot complex distributed systems across thousands of accelerators.

Qualifications

  • 5+ years of experience in software engineering with a focus on large-scale distributed systems or ML infrastructure.
  • Deep expertise in CUDA, NCCL, InfiniBand, and GPU cluster management.
  • Proficiency in Python, C++, and container orchestration tools like Kubernetes.

Responsibilities

  • Architect, build, and scale high-throughput, low-latency distributed training and inference systems.
  • Optimize cluster utilization, network topology, and storage subsystems for massive deep learning workloads.
  • Partner with research teams to remove infrastructure bottlenecks and accelerate model iteration speed.
  • Troubleshoot complex distributed systems issues across thousands of hardware accelerators.

Skills

CUDA
NCCL
InfiniBand
GPU cluster management
Python
C++
Kubernetes

Education

BS or MS in Computer Science or related technical discipline

Job description

OpenAI is seeking a Senior Infrastructure Engineer to help build and scale the compute infrastructure powering state-of-the-art generative AI models in San Francisco. You will work with massive clusters of GPUs to optimize training and inference performance.

You will collaborate with research teams to remove bottlenecks, optimize networks and storage, and troubleshoot complex distributed systems across thousands of accelerators.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior Infrastructure Engineer, AI
Senior Infrastructure Engineer, AI

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Infrastructure Engineer - GPU Compute
Senior AI Infrastructure Engineer - GPU Compute

Unchain Data • United States

On-site
USD 120,000 - 160,000
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000