Wayve is seeking a Cloud Site Reliability Engineer in Greater London to build and scale the reliability of its AI cloud platform. This founding role involves defining frameworks and operational standards while collaborating with teams to ensure system performance. Candidates should have strong Kubernetes and cloud systems support experience. The position also promotes a hybrid working model with in-office collaboration two days a week.
Qualifications
Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large‑scale cloud systems.
Strong Kubernetes experience, including operating production clusters.
Hands‑on experience running production workloads in AWS, GCP, or Azure.
Experience operating complex distributed systems in production.
Experience with large compute clusters and AI/ML workloads preferred.
Responsibilities
Own the reliability, availability, and performance of platform environments.
Participate in a 24/7 on‑call rotation for cloud incidents.
Design and operate monitoring and alerting systems.
Build automation for cluster operations and scaling tasks.
Skills
Kubernetes experience
Cloud systems support
Linux fundamentals
Scripting or systems language proficiency
Observability stacks design
Deep troubleshooting skills
Tools
AWS
GCP
Azure
Terraform
Datadog
Prometheus
Grafana
OpenTelemetry
Job description
Wayve is seeking a Cloud Site Reliability Engineer in Greater London to build and scale the reliability of its AI cloud platform. This founding role involves defining frameworks and operational standards while collaborating with teams to ensure system performance. Candidates should have strong Kubernetes and cloud systems support experience. The position also promotes a hybrid working model with in-office collaboration two days a week.