Software Engineer - Compute Infrastructure for Frontier AI

CV in

Northern (KY)

Hybrid

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a Software Engineer for Compute Infrastructure to build and optimize the platform that powers frontier AI. You will design, provision, and operate large-scale systems connecting GPUs, CPUs, networks, storage, and orchestration to run demanding workloads.

Ideal candidates have distributed systems experience, HPC familiarity, and a track record of reliability engineering, observability, and performance tuning at scale.

Qualifications

  • Strong software engineering skills and experience building, operating, or improving production infrastructure systems are essential.
  • Experience in distributed systems, operating systems, networking protocols, RDMA, NCCL or collective communication, storage, Kubernetes, scheduling, observability, reliability engineering, high‑performance computing, GPU infrastructure, CaaS, agent infrastructure, hardware‑aware performance optimization, benchmarking, developer experience, or infrastructure tooling is required.

Responsibilities

  • You will build and deeply optimize reliable system software for large-scale compute systems that run AI workloads.
  • You will design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health.
  • Profiling, benchmarking, and optimization of training workloads across compute, memory, storage, networking, NCCL and collective communication will form part of your daily work.

Skills

Distributed systems
High-performance computing
Observability
Reliability engineering

Tools

Kubernetes
NCCL/GPUs
RDMA
Observability tooling

Job description

OpenAI is seeking a Software Engineer for Compute Infrastructure to build and optimize the platform that powers frontier AI. You will design, provision, and operate large-scale systems connecting GPUs, CPUs, networks, storage, and orchestration to run demanding workloads.

Ideal candidates have distributed systems experience, HPC familiarity, and a track record of reliability engineering, observability, and performance tuning at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Linux Systems Engineer for Frontier Compute
Linux Systems Engineer for Frontier Compute

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Compute Infrastructure Engineer for Frontier AI
Compute Infrastructure Engineer for Frontier AI

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
Software Engineer, Hardware Health
Software Engineer, Hardware Health

Slope • San Francisco (CA)

On-site
USD 130,000 - 160,000
Software Engineer, Compute Foundations Systems
Software Engineer, Compute Foundations Systems

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

CV in • Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Power Management Engineer for Frontier Supercomputers
Power Management Engineer for Frontier Supercomputers

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 445,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Staff AI Infra Engineer - TPU & HPC Platforms
Staff AI Infra Engineer - TPU & HPC Platforms

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target 20%
Benefits package
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000