Compute Infrastructure Engineer for Frontier AI

OpenAI

California (MO)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking engineers to build and operate the compute platform that powers its research and products. You will work on large-scale, heterogeneous clusters spanning accelerators, CPUs, networks, and storage, with a focus on performance, reliability, and developer experience.

You may specialize near hardware or near users, contributing to CaaS, agent infrastructure, control/data planes, and tooling to make compute systems faster and easier to use.

Qualifications

  • Experience building, operating, or improving production infrastructure systems.
  • Strong background in distributed systems, networking, or HPC environments.
  • Familiarity with Kubernetes, NCCL, RDMA, and related tooling.

Responsibilities

  • Build and deeply optimize reliable system software for large-scale compute systems.
  • Design and operate infrastructure across accelerators, CPUs, NICs, switches, storage, and scheduling.
  • Profile, benchmark, and optimize training workloads across compute and networking bottlenecks.
  • Create hardware-aware automation for provisioning, firmware updates, and incident response.
  • Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools.
  • Turn operational lessons into better systems and clearer ownership boundaries across teams.
  • Collaborate across research, engineering, security, networking, hardware, and data center teams.

Skills

Strong software engineering
Distributed systems
Kubernetes
GPU infrastructure
CaaS
Observability
Profiling & benchmarking

Tools

NCCL
RDMA
GPU hardware tooling

Job description

OpenAI is seeking engineers to build and operate the compute platform that powers its research and products. You will work on large-scale, heterogeneous clusters spanning accelerators, CPUs, networks, and storage, with a focus on performance, reliability, and developer experience.

You may specialize near hardware or near users, contributing to CaaS, agent infrastructure, control/data planes, and tooling to make compute systems faster and easier to use.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Compute Infrastructure for Frontier AI
Software Engineer - Compute Infrastructure for Frontier AI

CV in • Northern (KY)

Hybrid
USD 180,000 - 240,000
Compute Infrastructure Engineer for Frontier AI
Compute Infrastructure Engineer for Frontier AI

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
Linux Systems Engineer for Frontier Compute
Linux Systems Engineer for Frontier Compute

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Linux Systems Engineer for Frontier Compute Platform
Linux Systems Engineer for Frontier Compute Platform

OpenAI • California (MO)

On-site
USD 140,000 - 210,000
Network Engineer: AI Infrastructure & High-Performance Networking
Network Engineer: AI Infrastructure & High-Performance Networking

OpenAI • California (MO)

On-site
USD 150,000 - 230,000
Frontier AI Infrastructure Engineer (Contract)
Frontier AI Infrastructure Engineer (Contract)

Mercor • Miami (FL)

On-site
USD 165,000 - 276,000
Network Engineering Lead for Frontier AI Compute
Network Engineering Lead for Frontier AI Compute

Socket.dev • New York (NY)

On-site
USD 180,000 - 280,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

CV in • Northern (KY)

Hybrid
USD 180,000 - 240,000
Accelerator Systems Software Engineer for AI Compute
Accelerator Systems Software Engineer for AI Compute

OpenAI • California (MO)

On-site
USD 180,000 - 300,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • California (MO)

On-site
USD 180,000 - 260,000