Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Job summary
Wayve, based in London, is seeking a Staff Cloud Site Reliability Engineer to establish and manage the reliability of their AI cloud platform. This role involves building the operational standards for their Model Development Platform and GPU Compute infrastructure, requiring extensive experience in large-scale cloud systems and Kubernetes. The position follows a hybrid working model with two days a week in the office, aiming to create a diverse and inclusive work environment.
Qualifications
Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
Strong Kubernetes experience, including operating production clusters.
Hands-on experience running production workloads in AWS, GCP, or Azure.
Responsibilities
Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments.
Participate in a 24/7 on-call rotation as first-line response for cloud and cluster-related incidents.
Design and operate monitoring, logging, tracing, and alerting systems.
Skills
SRE experience
Large-scale cloud systems
Kubernetes
AWS/GCP/Azure
Distributed systems
Linux fundamentals
Scripting languages (Python, Go, C++)
Observability tools
Tools
Terraform
Datadog
Prometheus
Grafana
Job description
Wayve, based in London, is seeking a Staff Cloud Site Reliability Engineer to establish and manage the reliability of their AI cloud platform. This role involves building the operational standards for their Model Development Platform and GPU Compute infrastructure, requiring extensive experience in large-scale cloud systems and Kubernetes. The position follows a hybrid working model with two days a week in the office, aiming to create a diverse and inclusive work environment.