Software Engineer, Compute Foundations

Linuxcareers

San Francisco, Northern (CA, KY)

Hybrid

USD 210,000 - 270,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OpenAI seeks engineers to design and operate Kubernetes-based controllers that coordinate GPU compute infrastructure across sites. You will define APIs, build provisioning services, and manage lifecycle operations from boot to decommissioning, ensuring reliability and scalability.

Ideal candidates have strong distributed-systems background, experience with Kubernetes reconciliation, and the ability to troubleshoot complex cross-boundary issues with staged rollouts across nodes and racks.

Qualifications

  • Strong software engineering fundamentals.
  • Experience designing production distributed systems or infrastructure services.
  • Experience with Kubernetes-based APIs and reconciliation patterns.
  • Ability to design reliable APIs and asynchronous workflows.
  • Capability to diagnose reliability and performance issues across borders.

Responsibilities

  • Design, build, and operate Kubernetes-based controllers and services across sites.
  • Define APIs and resource models for lifecycle operations across hardware providers.
  • Build provisioning/configuration services for network boot, firmware, and host config.
  • Develop lifecycle management for discovery, provisioning, upgrades, and recovery.
  • Design recovery mechanisms for concurrent changes and partial failures with staged rollouts.

Skills

Distributed systems
Kubernetes
APIs design
Troubleshooting

Job description

OpenAI's Compute Foundations team builds software that manages GPU compute infrastructure across data centers and sites, supporting model training and inference. In this role, you will design and operate Kubernetes-based distributed systems that provision, configure, and manage compute resources throughout their lifecycle, connecting global services with bare-metal systems management.

What You’ll Do
  • Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites and scale as GPU capacity grows
  • Define APIs and resource models that enable clients to request and track lifecycle operations across diverse hardware platforms and providers
  • Build provisioning and configuration services that coordinate network boot, hardware management interfaces, firmware deployment, operating-system images, drivers, and host configuration
  • Develop lifecycle management systems for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrated with health and validation systems
  • Design reliable reconciliation and recovery mechanisms for concurrent changes, interrupted operations, and partial failures with staged rollouts across nodes, racks, and clusters
What You Need
  • Strong software engineering fundamentals with experience designing, implementing, and owning production distributed systems or infrastructure services
  • Experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources
  • Understanding of bare-metal node provisioning from power-on to configured workload-ready state, with depth in areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management
  • Ability to design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries
  • Capability to diagnose reliability and performance problems across service, operating-system, and machine boundaries
Nice to Have
  • Experience building infrastructure control planes that coordinate operations across multiple sites or regions
  • Prior work with GPU or HPC infrastructure, including topology and shared dependencies across machines, racks, or clusters
  • Integration experience with multiple hardware platforms or infrastructure providers into a common service or resource model
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Compute Infrastructure Engineer
GPU Compute Infrastructure Engineer

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 490,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000