Founding Cloud SRE for AI Platform & GPU Clusters

Robotics Jobs UK

Greater London

Hybrid

GBP 90,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Wayve in London is hiring a founding Cloud Site Reliability Engineer to build and scale the reliability foundations of our AI cloud platform, including the Model Development Platform and the GPU Compute environment.

You will define frameworks, automation, and operational standards so our model development infrastructure and large compute clusters run predictably, efficiently and at scale, with 24/7 on-call responsibilities.

Qualifications

  • Proven SRE/Production eng for large-scale cloud systems.
  • Kubernetes in production.
  • Experience with AWS/GCP/Azure.
  • Experience with AI/ML training/inference workloads.
  • Strong Linux fundamentals and scripting.
  • Ability to design observability stacks.

Responsibilities

  • Own reliability of Model Dev Platform and GPU Compute.
  • Define SLOs/SLIs and budgets.
  • Improve capacity planning and scaling.
  • Lead incident response and postmortems.
  • Harden CI/CD and automation.
  • Build and operate monitoring systems.

Skills

Kubernetes
Cloud experience
Distributed systems
AI/ML workloads
Linux scripting
Troubleshooting
Observability tools
Communication

Tools

Datadog
Prometheus
Grafana
OpenTelemetry
CI/CD tooling

Job description

Wayve in London is hiring a founding Cloud Site Reliability Engineer to build and scale the reliability foundations of our AI cloud platform, including the Model Development Platform and the GPU Compute environment.

You will define frameworks, automation, and operational standards so our model development infrastructure and large compute clusters run predictably, efficiently and at scale, with 24/7 on-call responsibilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Staff SRE — AI Infrastructure & GPU Cloud
Founding Staff SRE — AI Infrastructure & GPU Cloud

Wayve • Greater London

Hybrid
GBP 110,000 - 150,000
Founding Cloud SRE: AI/ML Platform & GPU Compute
Founding Cloud SRE: AI/ML Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE - AI Platform & GPU Compute
Founding Cloud SRE - AI Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE — AI Platform & GPU Clusters (Hybrid)
Founding Cloud SRE — AI Platform & GPU Clusters (Hybrid)

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
Senior Cloud SRE - AI/ML Platform & GPU Compute
Senior Cloud SRE - AI/ML Platform & GPU Compute

Icehouseventures • Greater London

On-site
GBP 70,000 - 90,000
Senior Cloud SRE - AI/ML Platform & GPU Compute
Senior Cloud SRE - AI/ML Platform & GPU Compute

Wayve • Greater London

On-site
GBP 70,000 - 90,000
Staff SRE, AI Infrastructure
Staff SRE, AI Infrastructure

Wayve • Greater London

Hybrid
GBP 110,000 - 150,000
Staff SRE, AI Infrastructure
Staff SRE, AI Infrastructure

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Staff Cloud SRE - AI/ML Platform & GPU Compute
Staff Cloud SRE - AI/ML Platform & GPU Compute

Wayve • Greater London

On-site
GBP 70,000 - 90,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Robotics Jobs UK • Greater London

Hybrid
GBP 90,000 - 140,000