Senior Site Reliability Engineer — AI Inference Platform

Nebius

United States

Remote

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius is building a full-stack AI cloud platform addressing data and model training to production deployment. We seek an engineer to own the reliability, performance, and observability of the entire inference stack in a high-demand environment.

Responsibilities include designing telemetry pipelines, tuning autoscalers, and crafting infrastructure-as-code modules. You will collaborate with software engineers to deliver self-healing, scalable systems that meet aggressive cost and reliability

Qualifications

  • Deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and infrastructure-as-code practices.
  • You script in Python or Bash and understand alert design and SLOs for high-throughput APIs.
  • Experience running GPU-heavy workloads with accelerator stacks (e.g., vLLM, Triton, Ray) in production.

Responsibilities

  • Own reliability, performance, and observability of the inference stack.
  • Design telemetry pipelines (metrics, logs, traces) turning terabytes of data into actionable insight.
  • Tune Kubernetes autoscalers and craft Terraform modules to bake resilience into clusters; harden routing and retry logic.
  • Lead post-mortems and drive the culture to prevent recurrence.

Skills

Kubernetes
Prometheus
Grafana
Terraform
Python
Bash
SRE practices
Observability

Tools

Kubernetes
Prometheus
Grafana
Terraform
Python
Bash
vLLM / Triton / Ray

Job description

Nebius is building a full-stack AI cloud platform addressing data and model training to production deployment. We seek an engineer to own the reliability, performance, and observability of the entire inference stack in a high-demand environment.

Responsibilities include designing telemetry pipelines, tuning autoscalers, and crafting infrastructure-as-code modules. You will collaborate with software engineers to deliver self-healing, scalable systems that meet aggressive cost and reliability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Backend Engineer, AI Cloud Platform
Senior Backend Engineer, AI Cloud Platform

Nebius • Germany (OH)

On-site
USD 90,000 - 120,000
Competitive compensation
Career growth and learning opportunities
Flexibility and ownership
+3
Staff Software Engineer, AI Cloud Infrastructure
Staff Software Engineer, AI Cloud Infrastructure

Nebius • United States

On-site
USD 175,000 - 225,000
Health insurance
401(k) plan
Parental leave
+2
Global Network SRE for AI Cloud Infrastructure
Global Network SRE for AI Cloud Infrastructure

Socket.dev • United States

On-site
USD 180,000 - 224,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Network SRE: Reliability & Automation for Cloud Infra
Network SRE: Reliability & Automation for Cloud Infra

Nebius • United States

Remote
USD 140,000 - 210,000
Senior AI Inference SRE — Cloud & On‑Prem Reliability
Senior AI Inference SRE — Cloud & On‑Prem Reliability

SambaNova Systems • San Jose (CA)

On-site
USD 140,000 - 210,000
Competitive compensation
Equity and benefits
Senior Reliability Engineer, AI Inference at Scale
Senior Reliability Engineer, AI Inference at Scale

Cerebras • Raleigh (NC)

On-site
USD 170,000 - 250,000
Senior AI Cloud Infrastructure Engineer
Senior AI Cloud Infrastructure Engineer

Nebius • United States

Remote
USD 150,000 - 190,000
Competitive compensation
Career growth and learning
Flexible ownership and culture
+3
Senior Platform Testing Engineer – AI Inference Infra
Senior Platform Testing Engineer – AI Inference Infra

Baseten • United States

Remote
USD 140,000 - 210,000
Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

Socket.dev • United States

On-site
USD 180,000 - 224,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Site Reliability Engineer — AI Inference at Scale
Site Reliability Engineer — AI Inference at Scale

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1