Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 280,000 - 420,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity options
Health, vision, dental benefits
Unlimited PTO
Remote work options
Learning and development stipend

Job summary

OpenAI is seeking a Senior Infrastructure Engineer to design and scale ultra-high-performance distributed systems for training and inference of massive neural networks. You will optimize GPU cluster utilization, tackle networking bottlenecks, and collaborate with research teams to streamline experiments and model deployment.

The role requires deep expertise in HPC and distributed systems, with proficiency in C++, Python, CUDA, and Kubernetes, plus experience managing large-scale GPU clusters.

Qualifications

  • Requires a degree (B.S./M.S./Ph.D.) in Computer Science or related field with 5+ years of industry experience.
  • Deep expertise in distributed systems, HPC, and networking (InfiniBand, RoCE).
  • Proficiency in C++, Python, CUDA, and modern orchestration tools like Kubernetes.
  • Experience managing large-scale cloud or on-prem GPU clusters.

Responsibilities

  • Architect, build, and scale ultra-high-performance distributed systems for training and running massive neural networks.
  • Optimize GPU cluster utilization, networking bottlenecks, and memory hierarchies for large-scale training runs.
  • Collaborate with research teams to streamline experiment velocity and model deployment.
  • Troubleshoot complex infrastructure issues across thousands of accelerators.

Skills

Distributed systems
HPC
Networking
C++
Python
CUDA
Kubernetes

Education

BS/MS/PhD in Computer Science

Tools

InfiniBand
RoCE

Job description

OpenAI is seeking a Senior Infrastructure Engineer to design and scale ultra-high-performance distributed systems for training and inference of massive neural networks. You will optimize GPU cluster utilization, tackle networking bottlenecks, and collaborate with research teams to streamline experiments and model deployment.

The role requires deep expertise in HPC and distributed systems, with proficiency in C++, Python, CUDA, and Kubernetes, plus experience managing large-scale GPU clusters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
Senior AI Infra Engineer: Large-Scale GPU & HPC
Senior AI Infra Engineer: Large-Scale GPU & HPC

Anduril Industries • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Equity grants
Benefits package
Senior AI Infrastructure Engineer - Remote & Scaled GPU
Senior AI Infrastructure Engineer - Remote & Scaled GPU

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 160,000
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Senior AI Training Infra Engineer — GPU Clusters
Senior AI Training Infra Engineer — GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 300,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Senior Infrastructure Engineer, AI
Senior Infrastructure Engineer, AI

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000