Icehouseventures is seeking a Staff Cloud Site Reliability Engineer to shape the reliability of large-scale AI systems and GPU compute infrastructure. This founding role involves building and scaling reliability foundations for the AI cloud platform and ensuring cloud infrastructure resilience. Responsibilities include operationalizing SLOs, improving incident response, and creating automation for operations. The position offers a hybrid work model, encouraging collaboration in the London office while allowing remote work.
Qualifications
Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
Strong hands-on experience running production workloads in AWS, GCP, or Azure.
Deep troubleshooting skills across networking, storage, and distributed systems.
Responsibilities
Own the reliability and performance of the Model Dev Platform and GPU Compute environments.
Participate in a 24/7 on-call rotation as first-line response for incidents.
Design monitoring, logging, and alerting systems that enable rapid detection and recovery.
Skills
SRE or Production Engineer experience
Operating GPU-backed environments
MLOps experience
Strong Kubernetes
AWS/GCP/Azure experience
Distributed systems troubleshooting
Scripting or systems language proficiency
Designing observability stacks
Clear communication skills
Tools
Terraform
Datadog
Prometheus
Grafana
OpenTelemetry
Job description
Icehouseventures is seeking a Staff Cloud Site Reliability Engineer to shape the reliability of large-scale AI systems and GPU compute infrastructure. This founding role involves building and scaling reliability foundations for the AI cloud platform and ensuring cloud infrastructure resilience. Responsibilities include operationalizing SLOs, improving incident response, and creating automation for operations. The position offers a hybrid work model, encouraging collaboration in the London office while allowing remote work.