Senior SRE: AI Cloud Reliability & GPU Scale

Nebius Group

Greater London

On-site

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive pay
Career growth
Flexibility
Collaborative culture
Impactful AI projects
International team

Job summary

Nebius is seeking an experienced Site Reliability Engineer to own the reliability, performance, and observability of the full inference stack. You will design telemetry pipelines, tune Kubernetes autoscalers, and craft Terraform modules to ensure cost efficiency and resilience.

You’ll respond to incidents, drive post-morts, and collaborate with software engineers to turn reliability into a product feature. This role requires deep experience with Kubernetes, Prometheus, Grafana, Terraform, and

Qualifications

  • Deep fluency with Kubernetes, Prometheus, Grafana, and Terraform.
  • Proficiency in Python or Bash scripting.
  • Production experience with GPU-heavy workloads (vLLM, Triton, Ray).
  • Background in MLOps or model-hosting platforms; focus on reliability.

Responsibilities

  • Own reliability, performance, and observability of the inference stack.
  • Design telemetry pipelines—metrics, logs, traces—to drive insight.
  • Tune Kubernetes autoscalers to improve GPU efficiency and cost.
  • Craft Terraform modules to bake resilience into new clusters.
  • Harden request-routing and retry logic and drive post-mortems to prevent recurrence.

Skills

SRE practices
Distributed systems
Observability
Incident response

Tools

Kubernetes
Prometheus
Grafana
Terraform
Python
Bash
vLLM
Triton
Ray

Job description

Nebius is seeking an experienced Site Reliability Engineer to own the reliability, performance, and observability of the full inference stack. You will design telemetry pipelines, tune Kubernetes autoscalers, and craft Terraform modules to ensure cost efficiency and resilience.

You’ll respond to incidents, drive post-morts, and collaborate with software engineers to turn reliability into a product feature. This role requires deep experience with Kubernetes, Prometheus, Grafana, Terraform, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - AI Inference Platform, Scale & Reliability
Senior SRE - AI Inference Platform, Scale & Reliability

Nebius • Greater London

On-site
GBP 100,000 - 140,000
Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
+3
Senior SRE: Build Fault-Tolerant AI Cloud Infra
Senior SRE: Build Fault-Tolerant AI Cloud Infra

Nebius • Greater London

On-site
GBP 90,000 - 120,000
Competitive pay
Career growth
Flexible work
+3
Senior Cloud Reliability Engineer - Scale AI Infra & CI/CD
Senior Cloud Reliability Engineer - Scale AI Infra & CI/CD

Nebius • Greater London

On-site
GBP 90,000 - 120,000
Competitive compensation
Career growth
Learning opportunities
+5
Senior SRE for AI Platform & HPC
Senior SRE for AI Platform & HPC

Mistral Ai • York and North Yorkshire

Hybrid
GBP 70,000 - 110,000
Healthcare coverage
Relocation support
Wellness programs
Senior Cloud SRE: Scale AI Platform & Reliability
Senior Cloud SRE: Scale AI Platform & Reliability

Mistral AI • Greater London

On-site
GBP 75,000 - 110,000
Healthcare coverage
Relocation support
Retirement plans
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Nebius • Greater London

On-site
GBP 90,000 - 120,000
Competitive compensation
Career growth
Learning opportunities
+5
Senior SRE: Cloud Ops & CI/CD with Flexible Ownership
Senior SRE: Cloud Ops & CI/CD with Flexible Ownership

Nebius Group • Greater London

On-site
GBP 70,000 - 110,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Nebius Group • Greater London

On-site
GBP 70,000 - 110,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Nebius • Greater London

On-site
GBP 100,000 - 140,000
Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
+3
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Omilia Natural Language Solutions Ua Ltd • United Kingdom

On-site
GBP 85,000 - 120,000
Fixed compensation
Long-term employment with the working
Development in professional growth (ca
+1