Compute Infrastructure Engineer: Scale AI Compute

OpenAI

Greater London

On-site

GBP 171,258 - 301,563

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Remote work flexibility
Health insurance

Job summary

OpenAI is looking for engineers to build its compute platform tailored for advanced AI research and products in London. The role involves optimizing complex systems, debugging across various layers, and ensuring high performance.

Ideal candidates will have strong engineering judgment and expertise in distributed systems, high-performance computing, and network protocols. Join OpenAI to push the boundaries of AI technology.

Qualifications

  • Strong software engineering skills and experience building large-scale production systems.
  • Experience debugging across software, hardware, and networking layers.
  • Ability to optimize infrastructure for high-performance workloads.

Responsibilities

  • Build reliable system software for large-scale compute systems.
  • Optimize training workloads across various compute resources.
  • Collaborate across teams to enhance compute capacity and usability.

Skills

Distributed systems
High-performance computing
Reliability engineering
Kubernetes
Networking protocols
Debugging complex systems

Education

Bachelor's degree in Computer Science or related field

Tools

NCCL
RDMA
Monitoring tools

Job description

OpenAI is looking for engineers to build its compute platform tailored for advanced AI research and products in London. The role involves optimizing complex systems, debugging across various layers, and ensuring high performance.

Ideal candidates will have strong engineering judgment and expertise in distributed systems, high-performance computing, and network protocols. Join OpenAI to push the boundaries of AI technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Infrastructure Engineer — Scalable, Reliable Systems
Cloud Infrastructure Engineer — Scalable, Reliable Systems

OpenAI • Greater London

On-site
GBP 70,000 - 90,000
Infrastructure Engineering Manager, AI Platforms
Infrastructure Engineering Manager, AI Platforms

Scale AI • Greater London

On-site
GBP 110,000 - 180,000
GPU Infrastructure Engineer — Scale & Reliability for Production
GPU Infrastructure Engineer — Scale & Reliability for Production

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
GPU Infra Engineer — Scale, Automation & AI Compute
GPU Infra Engineer — Scale, Automation & AI Compute

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Cloud Infrastructure Engineer — Scale Distributed Systems
Cloud Infrastructure Engineer — Scale Distributed Systems

Slope • Greater London

On-site
GBP 60,000 - 80,000
Scale-Focused AI Infra & MLOps Engineer
Scale-Focused AI Infra & MLOps Engineer

EngineersOfAI • Greater London

On-site
GBP 90,000 - 120,000
Platform Engineer – Scale AI Infra & GPU Orchestration
Platform Engineer – Scale AI Infra & GPU Orchestration

Ineffable Intelligence • Greater London

On-site
GBP 70,000 - 110,000
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Senior AI Systems Engineer for Scalable GenAI Platform
Senior AI Systems Engineer for Scalable GenAI Platform

Nscale • Greater London

On-site
GBP 80,000 - 120,000
Software Engineer, Platform Systems
Software Engineer, Platform Systems

OpenAI • Greater London

On-site
GBP 60,000 - 80,000