Compute Foundations Engineer — Kubernetes & GPU Infra

OpenAI

San Francisco (CA)

On-site

USD 255,000 - 490,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OpenAI is seeking a Software Engineer for Compute Foundations in San Francisco to help manage GPU compute infrastructure across sites and data centers. You will design, build, and operate Kubernetes-based controllers and services that coordinate hardware, firmware, and software across clusters, while ensuring reliability as the fleet grows and evolves.

The role blends software architecture with on-prem and cloud hardware knowledge, offering a path to impact OpenAI’s research and products through

Qualifications

  • Experience designing, implementing, and owning production distributed systems or infrastructure services.
  • Experience developing infrastructure systems using Kubernetes APIs and reconciliation to manage resources.
  • Understanding how bare-metal nodes move from power-on to a configured, workload-ready state (PXE, DHCP/DNS, BMCs, firmware, Linux, drivers).

Responsibilities

  • Design, build, and operate Kubernetes-based controllers and distributed services coordinating infrastructure across sites.
  • Define APIs and resource models for lifecycle operations across hardware platforms and providers.
  • Build provisioning and configuration services coordinating network boot, hardware management interfaces, and deployment of firmware, OS images, drivers, and host config.
  • Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning with health validation integration.
  • Design reliable reconciliation and recovery handling concurrent changes and partial failures with staged rollouts.
  • Improve control-plane throughput and API latency while respecting site system and provider rate limits.
  • Build integrations bringing new sites and GPU hardware generations into the platform with cross-team collaboration.

Job description

OpenAI is seeking a Software Engineer for Compute Foundations in San Francisco to help manage GPU compute infrastructure across sites and data centers. You will design, build, and operate Kubernetes-based controllers and services that coordinate hardware, firmware, and software across clusters, while ensuring reliability as the fleet grows and evolves.

The role blends software architecture with on-prem and cloud hardware knowledge, offering a path to impact OpenAI’s research and products through

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

On-site
USD 210,000 - 270,000
GPU Compute Infrastructure Engineer
GPU Compute Infrastructure Engineer

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 490,000
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Compute Platform Engineer - GPU & Multi-Cloud Infra
Compute Platform Engineer - GPU & Multi-Cloud Infra

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Infra Engineer - GPU & Kubernetes
Senior AI Infra Engineer - GPU & Kubernetes

SB Telecom America Corp. • Sunnyvale (CA)

On-site
USD 150,000 - 250,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2