Senior SRE - AI Inference Platform, Scale & Reliability

Nebius

Greater London

On-site

GBP 100,000 - 140,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
Collaborative and innovative culture
Opportunity to work on impactful AI
International environment and talented

Job summary

Nebius in London is building a cutting‑edge AI cloud platform. This role owns the reliability, performance and observability of the entire inference stack, with hands‑on work on GPUs, Kubernetes and scalable cloud infrastructure.

You will implement telemetry pipelines (metrics, logs, traces), tune autoscalers with Kubernetes, craft Terraform modules, and harden routing and retries. You’ll drive post‑mortems to prevent recurrence while pushing for cost‑effective, high‑reliability operations that

Qualifications

  • Proven experience with Kubernetes-based infrastructure.
  • Deep knowledge of Prometheus, Grafana, and Terraform.
  • Experience with GPU workloads and ML deployment pipelines.
  • Background in MLOps or model-hosting platforms.

Responsibilities

  • Own reliability, performance and observability of the inference stack.
  • Design telemetry pipelines for metrics, logs and traces.
  • Tune autoscalers and Terraform modules to improve efficiency.
  • Lead post-mortems and drive remediation to prevent recurrence.

Skills

Python
Bash scripting
SRE fundamentals

Tools

Kubernetes
Prometheus
Grafana
Terraform

Job description

Nebius in London is building a cutting‑edge AI cloud platform. This role owns the reliability, performance and observability of the entire inference stack, with hands‑on work on GPUs, Kubernetes and scalable cloud infrastructure.

You will implement telemetry pipelines (metrics, logs, traces), tune autoscalers with Kubernetes, craft Terraform modules, and harden routing and retries. You’ll drive post‑mortems to prevent recurrence while pushing for cost‑effective, high‑reliability operations that

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Reliability Engineer - Scale AI Infra & CI/CD
Senior Cloud Reliability Engineer - Scale AI Infra & CI/CD

Nebius • Greater London

On-site
GBP 90,000 - 120,000
Competitive compensation
Career growth
Learning opportunities
+5
Senior SRE: Build Fault-Tolerant AI Cloud Infra
Senior SRE: Build Fault-Tolerant AI Cloud Infra

Nebius • Greater London

On-site
GBP 90,000 - 120,000
Competitive pay
Career growth
Flexible work
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Nebius • Greater London

On-site
GBP 90,000 - 120,000
Competitive compensation
Career growth
Learning opportunities
+5
Senior Backend Engineer - AI Cloud & Scalable Systems
Senior Backend Engineer - AI Cloud & Scalable Systems

Nebius • Greater London

On-site
GBP 110,000 - 160,000
Competitive pay
Career growth
Flexibility
+3
Senior Technical PM — AI Inference Platform (Remote/Hybrid)
Senior Technical PM — AI Inference Platform (Remote/Hybrid)

ApplyMint • United Kingdom

Hybrid
GBP 90,000 - 120,000
Competitive compensation
Career growth
Flexible work
+1
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Nebius • Greater London

On-site
GBP 100,000 - 140,000
Competitive compensation
Career growth and learning opportunity
Flexibility and ownership
+3
Senior Serverless AI Engineer - GPUs & Cloud, Hybrid
Senior Serverless AI Engineer - GPUs & Cloud, Hybrid

Nebius • Greater London

Hybrid
GBP 120,000 - 170,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior Cloud SRE for AI Platform — Reliability & Scale
Senior Cloud SRE for AI Platform — Reliability & Scale

Mistral • Greater London

On-site
GBP 90,000 - 140,000
Healthcare coverage
Parental leave
Retirement plans
+3
Senior SRE: Cloud & Edge Reliability Engineer
Senior SRE: Cloud & Edge Reliability Engineer

Atarus • Greater London

On-site
GBP 90,000 - 150,000
Senior Cloud SRE: Scale AI Platform & Reliability
Senior Cloud SRE: Scale AI Platform & Reliability

Mistral AI • Greater London

On-site
GBP 75,000 - 110,000
Healthcare coverage
Relocation support
Retirement plans
+3