GPU Infra Engineer — Scale, Automation & AI Compute

OpenAI

Greater London

On-site

GBP 120,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OpenAI is seeking a Software Engineer to design and operate large-scale GPU infrastructure powering ChatGPT workloads. You will build systems for fleet health, capacity planning, and automation, collaborating with infrastructure, research, and product teams to improve reliability and compute utilization.

This role suits engineers who enjoy solving operational challenges, building internal platforms, and advancing infrastructure that scales with frontier AI in a fast-moving environment.

Qualifications

  • Five or more years of software engineering experience building production infrastructure.
  • Experience designing and operating highly available distributed systems.
  • Experience with GPU infrastructure, HPC, ML infrastructure, or large-scale compute platforms.
  • Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
  • Excellent debugging, systems design, and operational problem-solving skills.

Responsibilities

  • Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.
  • Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.
  • Improve observability, reliability, and operational efficiency across thousands of GPUs.
  • Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
  • Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.
  • Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform.
  • Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.

Skills

Go
Python
C++
Rust
Kubernetes
Linux
Observability
Distributed systems

Tools

Terraform
Docker
Prometheus

Job description

OpenAI is seeking a Software Engineer to design and operate large-scale GPU infrastructure powering ChatGPT workloads. You will build systems for fleet health, capacity planning, and automation, collaborating with infrastructure, research, and product teams to improve reliability and compute utilization.

This role suits engineers who enjoy solving operational challenges, building internal platforms, and advancing infrastructure that scales with frontier AI in a fast-moving environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer — Scale & Reliability for Production
GPU Infrastructure Engineer — Scale & Reliability for Production

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
Compute Infrastructure Engineer: Scale AI Compute
Compute Infrastructure Engineer: Scale AI Compute

OpenAI • Greater London

On-site
GBP 171,000 - 302,000
Competitive salary
Remote work flexibility
Health insurance
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Senior Backend Engineer, Scalable ChatGPT Infra
Senior Backend Engineer, Scalable ChatGPT Infra

OpenAI • Greater London

On-site
GBP 191,000 - 305,000
Cloud Infrastructure Engineer — Scalable, Reliable Systems
Cloud Infrastructure Engineer — Scalable, Reliable Systems

OpenAI • Greater London

On-site
GBP 70,000 - 90,000
Backend Engineer — Scalable ChatGPT Infrastructure
Backend Engineer — Scalable ChatGPT Infrastructure

Ritual Ads ® • Greater London

On-site
GBP 90,000 - 125,000
GenAI Infra Engineer: Real-Time GPU Serving & Open-Weights
GenAI Infra Engineer: Real-Time GPU Serving & Open-Weights

deliveroo • Greater London

On-site
GBP 90,000 - 150,000
Software Engineer, Model Deployment- ChatGPT Engineering
Software Engineer, Model Deployment- ChatGPT Engineering

OpenAI • Greater London

On-site
GBP 194,000 - 280,000