Founding Cloud SRE - AI Platform & GPU Compute

Wayve

Greater London

Hybrid

GBP 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Wayve is seeking a Cloud Site Reliability Engineer in Greater London to build and scale the reliability of its AI cloud platform. This founding role involves defining frameworks and operational standards while collaborating with teams to ensure system performance. Candidates should have strong Kubernetes and cloud systems support experience. The position also promotes a hybrid working model with in-office collaboration two days a week.

Qualifications

  • Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large‑scale cloud systems.
  • Strong Kubernetes experience, including operating production clusters.
  • Hands‑on experience running production workloads in AWS, GCP, or Azure.
  • Experience operating complex distributed systems in production.
  • Experience with large compute clusters and AI/ML workloads preferred.

Responsibilities

  • Own the reliability, availability, and performance of platform environments.
  • Participate in a 24/7 on‑call rotation for cloud incidents.
  • Design and operate monitoring and alerting systems.
  • Build automation for cluster operations and scaling tasks.

Skills

Kubernetes experience
Cloud systems support
Linux fundamentals
Scripting or systems language proficiency
Observability stacks design
Deep troubleshooting skills

Tools

AWS
GCP
Azure
Terraform
Datadog
Prometheus
Grafana
OpenTelemetry

Job description

Wayve is seeking a Cloud Site Reliability Engineer in Greater London to build and scale the reliability of its AI cloud platform. This founding role involves defining frameworks and operational standards while collaborating with teams to ensure system performance. Candidates should have strong Kubernetes and cloud systems support experience. The position also promotes a hybrid working model with in-office collaboration two days a week.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Cloud SRE: AI/ML Platform & GPU Compute
Founding Cloud SRE: AI/ML Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE — AI Platform & GPU Clusters (Hybrid)
Founding Cloud SRE — AI Platform & GPU Clusters (Hybrid)

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE for AI Platform & GPU Clusters
Founding Cloud SRE for AI Platform & GPU Clusters

Robotics Jobs UK • Greater London

Hybrid
GBP 90,000 - 140,000
Founding Staff SRE — AI Infrastructure & GPU Cloud
Founding Staff SRE — AI Infrastructure & GPU Cloud

Wayve • Greater London

Hybrid
GBP 110,000 - 150,000
Senior Cloud SRE - AI/ML Platform & GPU Compute
Senior Cloud SRE - AI/ML Platform & GPU Compute

Icehouseventures • Greater London

On-site
GBP 70,000 - 90,000
Senior Cloud SRE - AI/ML Platform & GPU Compute
Senior Cloud SRE - AI/ML Platform & GPU Compute

Wayve • Greater London

On-site
GBP 70,000 - 90,000
Founding Cloud SRE — AI/ML Platform & GPU Compute
Founding Cloud SRE — AI/ML Platform & GPU Compute

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
Staff SRE, AI Infrastructure
Staff SRE, AI Infrastructure

Wayve • Greater London

Hybrid
GBP 110,000 - 150,000
Staff Cloud SRE - AI/ML Platform & GPU Compute
Staff Cloud SRE - AI/ML Platform & GPU Compute

Wayve • Greater London

On-site
GBP 70,000 - 90,000
Staff SRE, AI Infrastructure
Staff SRE, AI Infrastructure

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000