Senior SRE - Compute Nodes (Linux & Virtualization)

Nebius

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nebius is seeking a Senior Site Reliability Engineer to join the Compute Node team. The role focuses on Linux systems engineering, virtualization, and operational reliability across multi-region compute nodes that run and manage VMs.

You will work close to the OS and hypervisor, shaping reliability and observability from the ground up while leading incident response and postmortems to improve long-term system health.

Qualifications

  • Strong Linux user and kernel space knowledge.
  • Experience with virtualization tech in production.
  • Experience designing observability stacks and reliability practices.

Responsibilities

  • Ensure reliability, availability and performance of compute nodes running VMs
  • Analyze and debug Linux systems across user space and kernel space
  • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
  • Work hands-on with virtualization and containerization using QEMU/KVM and Linux-native tech
  • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
  • Lead incident response, root-cause analysis, and postmortems to drive reliability improvements
  • Collaborate with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability

Skills

Linux expertise
Virtualization
Containerization
Debugging skills
SRE mindset

Tools

QEMU/KVM

Job description

Nebius is seeking a Senior Site Reliability Engineer to join the Compute Node team. The role focuses on Linux systems engineering, virtualization, and operational reliability across multi-region compute nodes that run and manage VMs.

You will work close to the OS and hypervisor, shaping reliability and observability from the ground up while leading incident response and postmortems to improve long-term system health.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network SRE: Reliability & Automation for Cloud Infra
Network SRE: Reliability & Automation for Cloud Infra

Nebius • United States

Remote
USD 140,000 - 210,000
Remote SRE, Hardware Infra for AI Cloud
Remote SRE, Hardware Infra for AI Cloud

Nebius • United States

On-site
USD 130,000 - 180,000
100% company-paid medical, dental, and vision insurance
401(k) plan with company match
20 weeks paid parental leave for primary caregivers
+2
Global Network SRE for AI Cloud Infrastructure
Global Network SRE for AI Cloud Infrastructure

Socket.dev • United States

On-site
USD 180,000 - 224,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops

Nebius • United States

Remote
USD 120,000 - 170,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Remote SRE — AI Cloud Hardware Infra
Remote SRE — AI Cloud Hardware Infra

Nebius • United States

On-site
USD 130,000 - 180,000
Health insurance
401(k) plan
Parental leave
+2
Senior SRE - Kubernetes, GitOps & Observability (NYC)
Senior SRE - Kubernetes, GitOps & Observability (NYC)

Nebius B.V. • New York (NY)

On-site
USD 147,000 - 224,000
Health insurance
401(k)
Parental leave
+1
Senior Compute Platform SRE - Bare-Metal & Kubernetes Remote
Senior Compute Platform SRE - Bare-Metal & Kubernetes Remote

LaSalle Network • Chicago (IL)

On-site
USD 110,000 - 124,000
Benefits eligible
Senior Staff SRE — Global Infra, Automation & Observability
Senior Staff SRE — Global Infra, Automation & Observability

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 322,000
Senior SRE: Global Infra & Automation (Equity)
Senior SRE: Global Infra & Automation (Equity)

Socket.dev • California (MO)

Hybrid
USD 200,000 - 322,000
Equity
Benefits