GPU Infrastructure Engineer — Scale & Reliability for Production

Triwill Group

Greater London

Hybrid

GBP 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI in London seeks a software engineer to build and operate large-scale production systems powering ChatGPT. You’ll develop tooling for fleet health, capacity planning, automation, and incident response, collaborating with infra, research, and product teams to improve reliability and compute utilization.

The role suits engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale.

Qualifications

  • Five or more years of software engineering experience building production infrastructure.
  • Strong programming skills in Go, Python, C++, Rust, or a comparable language.
  • Experience designing or operating highly available distributed systems.
  • Experience with GPU infrastructure, high-performance computing, ML infra, or large-scale compute platforms.

Responsibilities

  • Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference.
  • Develop systems for capacity planning, fleet health monitoring, and resource utilization.
  • Automate operational workflows, including incident detection, diagnosis, and response.
  • Identify and address bottlenecks affecting fleet reliability, scalability, and performance.
  • Partner with infrastructure, research, and product engineering teams to improve the compute platform.

Skills

Go
Python
C++
Rust

Job description

OpenAI in London seeks a software engineer to build and operate large-scale production systems powering ChatGPT. You’ll develop tooling for fleet health, capacity planning, automation, and incident response, collaborating with infra, research, and product teams to improve reliability and compute utilization.

The role suits engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infra Engineer — Scale, Automation & AI Compute
GPU Infra Engineer — Scale, Automation & AI Compute

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
GPU Infrastructure Engineer for Scalable AI Inference
GPU Infrastructure Engineer for Scalable AI Inference

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Backend Engineer, Scalable ChatGPT Infra
Senior Backend Engineer, Scalable ChatGPT Infra

OpenAI • Greater London

On-site
GBP 191,000 - 305,000
Backend Engineer — Scalable ChatGPT Infrastructure
Backend Engineer — Scalable ChatGPT Infrastructure

Ritual Ads ® • Greater London

On-site
GBP 90,000 - 125,000
Cloud Infrastructure Engineer — Scalable, Reliable Systems
Cloud Infrastructure Engineer — Scalable, Reliable Systems

OpenAI • Greater London

On-site
GBP 70,000 - 90,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Cloud Infrastructure Engineer — Scale Distributed Systems
Cloud Infrastructure Engineer — Scale Distributed Systems

Slope • Greater London

On-site
GBP 60,000 - 80,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
Software Engineer, ChatGPT Infrastructure
Software Engineer, ChatGPT Infrastructure

OpenAI • Greater London

On-site
GBP 191,000 - 305,000