GPU Infrastructure Engineer for Scalable AI Inference

AI Startups UK

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a software engineer to build and operate large-scale production systems that power ChatGPT's GPU fleet. You will develop tooling for fleet health, capacity planning, automation, and incident response, working with infrastructure, research, and product teams.

Candidates should have 5+ years in software engineering, strong skills in Go, Python, C++, or Rust, and experience with distributed systems and GPU infrastructure.

Qualifications

  • Five or more years of software engineering experience building production infrastructure.
  • Strong programming skills in Go, Python, C++, or Rust.
  • Experience designing or operating highly available distributed systems.
  • Experience with GPU infrastructure or large-scale compute platforms.
  • Strong debugging and operational problem-solving skills.
  • Excellent cross-team communication.

Responsibilities

  • Build software and internal tools to manage large-scale GPU infrastructure powering ChatGPT inference.
  • Develop systems for capacity planning, fleet health monitoring, and resource utilization.
  • Automate operational workflows, including incident detection, diagnosis, and response.
  • Identify bottlenecks affecting fleet reliability, scalability, and performance.
  • Partner with infrastructure, research, and product engineering teams to improve the compute platform.

Skills

Go
Python
C++
Rust
Distributed systems
Production infrastructure
Debugging
Communication

Tools

GPU infrastructure
Cluster orchestration

Job description

OpenAI is seeking a software engineer to build and operate large-scale production systems that power ChatGPT's GPU fleet. You will develop tooling for fleet health, capacity planning, automation, and incident response, working with infrastructure, research, and product teams.

Candidates should have 5+ years in software engineering, strong skills in Go, Python, C++, or Rust, and experience with distributed systems and GPU infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infra Engineer — Scale, Automation & AI Compute
GPU Infra Engineer — Scale, Automation & AI Compute

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
GPU Infrastructure Engineer — Scale & Reliability for Production
GPU Infrastructure Engineer — Scale & Reliability for Production

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
AI Infrastructure Engineer: GPU & Performance Optimisation
AI Infrastructure Engineer: GPU & Performance Optimisation

twentyAI • Greater London

On-site
GBP 90,000 - 140,000
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
Site Reliability Engineer, GPUs in AI
Site Reliability Engineer, GPUs in AI

Radley James • Greater London

On-site
GBP 20,000 - 40,000
GenAI Infra Engineer: Real-Time GPU Serving & Fine-Tuning
GenAI Infra Engineer: Real-Time GPU Serving & Fine-Tuning

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Senior Backend Engineer, Scalable ChatGPT Infra
Senior Backend Engineer, Scalable ChatGPT Infra

OpenAI • Greater London

On-site
GBP 191,000 - 305,000