Founding Staff SRE — AI Infrastructure & GPU Cloud

Wayve

Greater London

Hybrid

GBP 110,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Wayve, a London-based AI robotics company, is hiring a Staff Cloud Site Reliability Engineer to help found and scale the reliability foundations of our AI cloud platform.

You will own the Model Dev Platform and GPU Compute environments, define production standards, and lead incident response, observability, and automation efforts to ensure resilient, high-performance systems across multi-tenant clusters.

Qualifications

  • Proven experience in an SRE/Production Engineer/Cloud Reliability role.
  • Experience operating GPU-backed environments or large-scale ML infrastructure.
  • Experience running model training or inference pipelines in production (MLOps).
  • Strong Kubernetes experience, including operating production clusters.
  • Hands-on experience running production workloads in AWS, GCP, or Azure.
  • Experience operating complex distributed systems in production, ideally with compute-heavy workloads.
  • Experience with large compute clusters and AI/ML workloads is preferred.
  • Strong Linux fundamentals and scripting or systems language with automation bias.
  • Experience designing observability stacks (Datadog, Prometheus, Grafana, OpenTelemetry).
  • Clear communication: incidents, postmortems, reliability improvements.

Responsibilities

  • Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments.
  • Define and operationalise SLOs, SLIs, and error budgets across platform services.
  • Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters.
  • Partner with ML, platform, and software teams to establish production readiness standards.
  • Participate in a 24/7 on-call rotation for cloud and cluster incidents.
  • Lead incident triage, escalation, and root-cause analysis; implement durable improvements.
  • Design and operate monitoring, logging, tracing, and alerting systems for rapid recovery.
  • Build dashboards reflecting real user-centric platform health and reliability.

Skills

SRE experience
Production engineering
Cloud reliability
Linux fundamentals
Troubleshooting skills

Tools

Kubernetes
AWS
GCP
Azure
Datadog
Prometheus
Grafana
OpenTelemetry

Job description

Wayve, a London-based AI robotics company, is hiring a Staff Cloud Site Reliability Engineer to help found and scale the reliability foundations of our AI cloud platform.

You will own the Model Dev Platform and GPU Compute environments, define production standards, and lead incident response, observability, and automation efforts to ensure resilient, high-performance systems across multi-tenant clusters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Cloud SRE: AI/ML Platform & GPU Compute
Founding Cloud SRE: AI/ML Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE - AI Platform & GPU Compute
Founding Cloud SRE - AI Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Staff SRE, AI Infrastructure
Staff SRE, AI Infrastructure

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Staff SRE, AI Infrastructure
Staff SRE, AI Infrastructure

Wayve • Greater London

Hybrid
GBP 110,000 - 150,000
Senior Cloud SRE - AI/ML Platform & GPU Compute
Senior Cloud SRE - AI/ML Platform & GPU Compute

Wayve • Greater London

On-site
GBP 70,000 - 90,000
Senior Platform Engineer — Cloud Infra, CI/CD & DX
Senior Platform Engineer — Cloud Infra, CI/CD & DX

Robotics Jobs UK • Greater London

Hybrid
GBP 70,000 - 110,000
Senior Software Engineer, AI Portal — Product & Cloud
Senior Software Engineer, AI Portal — Product & Cloud

Lindus Health • Greater London

Hybrid
GBP 90,000 - 130,000
Hybrid work policy
Software Engineer, AI Platforms & Libraries
Software Engineer, AI Platforms & Libraries

Wayve • Greater London

On-site
GBP 90,000 - 135,000
Full-Stack Platform Engineer, Model Development
Full-Stack Platform Engineer, Model Development

EngineersOfAI • Greater London

Hybrid
GBP 90,000 - 120,000
Senior Data & MLOps Engineer - AI Reliability Platform
Senior Data & MLOps Engineer - AI Reliability Platform

CoreWeave • Greater London

On-site
GBP 120,000 - 190,000
Medical Insurance
Dental Insurance
Pension
+5