Platform Systems Engineer: Scalable AI Infra & Observability

OpenAI

Greater London

On-site

GBP 60,000 - 80,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A global AI research company in London is seeking a Software Engineer for Platform Systems to enhance large-scale AI training infrastructure. Key responsibilities include designing failure detection systems, improving observability, and collaborating with various teams to ensure the platform's reliability. Ideal candidates should have experience in distributed systems, performance optimization, and debugging complex issues. Join us in shaping the future of AI technology with state-of-the-art engineering solutions.

Qualifications

  • Experience with large-scale distributed systems and their performance.
  • Experience writing low-level software.
  • Understanding of hardware, operating systems, and networking.

Responsibilities

  • Design and build failure detection and tracing systems for AI training jobs.
  • Develop tools to identify issues in large-scale systems.
  • Improve observability and reliability of training infrastructure.

Skills

Performance analysis
Debugging
Systems engineering
Distributed systems
Observability

Job description

A global AI research company in London is seeking a Software Engineer for Platform Systems to enhance large-scale AI training infrastructure. Key responsibilities include designing failure detection systems, improving observability, and collaborating with various teams to ensure the platform's reliability. Ideal candidates should have experience in distributed systems, performance optimization, and debugging complex issues. Join us in shaping the future of AI technology with state-of-the-art engineering solutions.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Systems Engineer: Scalable AI Training & Telemetry
Platform Systems Engineer: Scalable AI Training & Telemetry

OpenAI • City Of London

On-site
GBP 70,000 - 90,000
Senior ML Platform Engineer - Build Scalable AI Infra
Senior ML Platform Engineer - Build Scalable AI Infra

Synthesia • Greater London

On-site
GBP 70,000 - 100,000
Senior Platform Engineer, Data & AI Platform
Senior Platform Engineer, Data & AI Platform

albelli • London

On-site
GBP 70,000 - 90,000
Software Engineer, Platform Systems
Software Engineer, Platform Systems

OpenAI • Greater London

On-site
GBP 60,000 - 80,000
Senior AI Systems Engineer for Scalable GenAI Platform
Senior AI Systems Engineer for Scalable GenAI Platform

Nscale • Greater London

On-site
GBP 80,000 - 120,000
Infrastructure Engineering Manager, AI Platforms
Infrastructure Engineering Manager, AI Platforms

Scale AI • Greater London

On-site
GBP 110,000 - 180,000
Senior Backend Engineer - AI Platform & Scalable Systems
Senior Backend Engineer - AI Platform & Scalable Systems

Gradient Labs Limited • Greater London

Hybrid
GBP 70,000 - 90,000
Platform Engineer – Scale AI Infra & GPU Orchestration
Platform Engineer – Scale AI Infra & GPU Orchestration

Ineffable Intelligence • Greater London

On-site
GBP 70,000 - 110,000
AI Software Engineer — Build Scalable AI in the Cloud
AI Software Engineer — Build Scalable AI in the Cloud

iManage • City Of London

On-site
GBP 90,000 - 140,000
Flexible work hours
Market-leading salary
Annual performance-based bonus
+4
Senior Platform Engineer: AI-Driven Infra & Cloud
Senior Platform Engineer: AI-Driven Infra & Cloud

Sr2 Rec Ltd • Greater London

On-site
GBP 100,000 - 110,000
Meaningful equity
Excellent benefits