Founding Cloud SRE — AI/ML Platform & GPU Compute

Icehouseventures

Greater London

Hybrid

GBP 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Icehouseventures is seeking a Staff Cloud Site Reliability Engineer to shape the reliability of large-scale AI systems and GPU compute infrastructure. This founding role involves building and scaling reliability foundations for the AI cloud platform and ensuring cloud infrastructure resilience. Responsibilities include operationalizing SLOs, improving incident response, and creating automation for operations. The position offers a hybrid work model, encouraging collaboration in the London office while allowing remote work.

Qualifications

  • Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
  • Strong hands-on experience running production workloads in AWS, GCP, or Azure.
  • Deep troubleshooting skills across networking, storage, and distributed systems.

Responsibilities

  • Own the reliability and performance of the Model Dev Platform and GPU Compute environments.
  • Participate in a 24/7 on-call rotation as first-line response for incidents.
  • Design monitoring, logging, and alerting systems that enable rapid detection and recovery.

Skills

SRE or Production Engineer experience
Operating GPU-backed environments
MLOps experience
Strong Kubernetes
AWS/GCP/Azure experience
Distributed systems troubleshooting
Scripting or systems language proficiency
Designing observability stacks
Clear communication skills

Tools

Terraform
Datadog
Prometheus
Grafana
OpenTelemetry

Job description

Icehouseventures is seeking a Staff Cloud Site Reliability Engineer to shape the reliability of large-scale AI systems and GPU compute infrastructure. This founding role involves building and scaling reliability foundations for the AI cloud platform and ensuring cloud infrastructure resilience. Responsibilities include operationalizing SLOs, improving incident response, and creating automation for operations. The position offers a hybrid work model, encouraging collaboration in the London office while allowing remote work.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Cloud SRE: AI/ML Platform & GPU Compute
Founding Cloud SRE: AI/ML Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE - AI Platform & GPU Compute
Founding Cloud SRE - AI Platform & GPU Compute

Wayve • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Cloud SRE for AI Platform & GPU Clusters
Founding Cloud SRE for AI Platform & GPU Clusters

Robotics Jobs UK • Greater London

Hybrid
GBP 90,000 - 140,000
Founding Cloud SRE — AI Platform & GPU Clusters (Hybrid)
Founding Cloud SRE — AI Platform & GPU Clusters (Hybrid)

Icehouseventures • Greater London

Hybrid
GBP 70,000 - 90,000
Founding Staff SRE — AI Infrastructure & GPU Cloud
Founding Staff SRE — AI Infrastructure & GPU Cloud

Wayve • Greater London

Hybrid
GBP 110,000 - 150,000
Senior Cloud SRE: Scale AI Platform & Reliability
Senior Cloud SRE: Scale AI Platform & Reliability

Mistral AI • Greater London

On-site
GBP 75,000 - 110,000
Healthcare coverage
Relocation support
Retirement plans
+3
Senior Cloud SRE for AI Platform — Reliability & Scale
Senior Cloud SRE for AI Platform — Reliability & Scale

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare coverage
Parental leave
Retirement plans
+3
Remote Cloud SRE: Scalable Systems & Automation
Remote Cloud SRE: Scalable Systems & Automation

Intapp, Inc. • Greater London

Remote
GBP 75,000 - 95,000
Professional development opportunities
Flexible work environment
Comprehensive wellness programs
Founding Applied AI SRE Engineer for Platform Reliability
Founding Applied AI SRE Engineer for Platform Reliability

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare
Relocation
Wellness program
+1
Senior AI Reliability Engineer | Large-Scale Model Serving
Senior AI Reliability Engineer | Large-Scale Model Serving

Anthropic • Greater London

Hybrid
GBP 80,000 - 120,000
Health insurance
Dental insurance
Vision insurance
+1