Compute Infrastructure Engineer for Frontier AI

openai

General Trias

On-site

PHP 11,243,000 - 18,738,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

OpenAI is seeking engineers to build the compute platform behind its research and products. You may work near hardware or near users, on CaaS and agent infrastructure, or on the control and data planes in between.

You could help bring new supercomputing capacity online, optimize training workloads, improve NCCL and collective communication, and design abstractions that make heterogeneous clusters feel like one coherent platform.

Qualifications

  • Strong software engineering skills with production infrastructure experience.
  • Experience in distributed systems, HPC, GPU infrastructure, or related areas.
  • Ability to debug complex system behavior across software, hardware, networking, and workload layers.

Responsibilities

  • Build and optimize reliable system software for large-scale compute systems.
  • Design and operate infrastructure across accelerators, CPUs, NICs, networking, storage, data centers, cluster orchestration, scheduling, and fleet health.
  • Profile, benchmark, and optimize training workloads and scheduling bottlenecks across compute, memory, storage, networking, NCCL.
  • Create hardware-aware automation for provisioning, firmware and driver upgrades, incident response, and day-to-day operations.
  • Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools for researchers, engineers, and operators.

Skills

Distributed systems
Operating systems
Networking protocols
High-performance computing
GPU infrastructure
CaaS
Agent infrastructure
Scheduling
Reliability engineering
Developer platforms
Instrumentation & observability

Tools

Kubernetes
NCCL
RDMA
GPU tooling
Observability tooling

Job description

OpenAI is seeking engineers to build the compute platform behind its research and products. You may work near hardware or near users, on CaaS and agent infrastructure, or on the control and data planes in between.

You could help bring new supercomputing capacity online, optimize training workloads, improve NCCL and collective communication, and design abstractions that make heterogeneous clusters feel like one coherent platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

openai • General Trias

On-site
PHP 11,243,000 - 18,738,000
Agent Infrastructure Engineer - Scale AI Training (Hybrid)
Agent Infrastructure Engineer - Scale AI Training (Hybrid)

openai • General Trias

Hybrid
PHP 13,117,000 - 16,240,000
Relocation assistance
Hybrid work model
GenAI Full-Stack Engineer: Build & Deploy Frontier Models
GenAI Full-Stack Engineer: Build & Deploy Frontier Models

HumanitApp • Hinoba-an

On-site
PHP 2,159,000 - 2,591,000
Software Engineer, Agent Infrastructure
Software Engineer, Agent Infrastructure

openai • General Trias

Hybrid
PHP 13,117,000 - 16,240,000
Relocation assistance
Hybrid work model
AI Infra Engineer: Scalable Agent Systems
AI Infra Engineer: Scalable Agent Systems

Proximal • Hinoba-an

On-site
PHP 600,000 - 1,000,000
Infrastructure Engineer, Codex Core Agents
Infrastructure Engineer, Codex Core Agents

openai • General Trias

On-site
PHP 7,495,000 - 11,243,000
Platform Engineer for AI-Driven Physics Infrastructure
Platform Engineer for AI-Driven Physics Infrastructure

Embedded Shishya • Boston

Hybrid
PHP 7,385,000 - 11,077,000
Software Engineer
Software Engineer

Proximal • Hinoba-an

On-site
PHP 600,000 - 1,000,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Fireworks • San Mateo

On-site
PHP 1,800,000 - 3,000,000
Staff AI Platform Engineer: Scalable AI Infrastructure
Staff AI Platform Engineer: Scalable AI Infrastructure

CodeRound • Hinoba-an

On-site
PHP 1,500,000 - 2,400,000