Site Reliability Engineer in Network Infrastructure

Nebius Group

United States

Remote

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Nebius is seeking a Site Reliability Engineer to own the Network infrastructure, defining reliability targets and building tooling to scale reliably. This engineering-first role focuses on measurable SLIs/SLOs, robust change workflows, and strong operability across the global network.

You will collaborate with network engineers and platform teams to embed observability, automate incident response, and ensure rapid recovery from issues as Nebius expands in North America and beyond.

Responsibilities

  • Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense)
  • Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards
  • Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting)
  • Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents
  • Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes
  • Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast

Job description

About Nebius : Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The Role

We’re looking for a Site Reliability Engineer to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly. Your responsibilities will include:

  • Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense)
  • Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards
  • Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting)
  • Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents
  • Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes
  • Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast

We expect you to have: Strong p

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer in Network Infrastructure
Software Engineer in Network Infrastructure

Nebius Group • United States

Remote
USD 120,000 - 150,000
Network SRE: Reliability, Automation & Observability
Network SRE: Reliability, Automation & Observability

Nebius Group • United States

Remote
USD 120,000 - 180,000
Network SRE: Reliability & Automation for Cloud Infra
Network SRE: Reliability & Automation for Cloud Infra

Nebius • United States

Remote
USD 140,000 - 210,000
Senior Software Engineer, Observability
Senior Software Engineer, Observability

Nebius • United States

On-site
USD 130,000 - 170,000
100% company-paid health insurance
401(k) match
Parental leave
+1
Senior Software Engineer, Network Automation
Senior Software Engineer, Network Automation

Nebius Group • United States

Remote
USD 120,000 - 150,000
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops

Nebius • United States

Remote
USD 120,000 - 170,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Site Reliability Engineer (SRE) - Early Talent
Site Reliability Engineer (SRE) - Early Talent

Nebius Group • City of Amsterdam (NY)

On-site
USD 15,000 - 23,000
Mentorship
Hands-on production experience
Growth opportunities
+2
Data Engineer
Data Engineer

Meyandy LLC • Northern (KY)

On-site
USD 120,000 - 190,000
Forward Deployed Engineer - Physical AI
Forward Deployed Engineer - Physical AI

Nebius • United States

Remote
USD 150,000 - 210,000
Senior Network Infrastructure Software Engineer
Senior Network Infrastructure Software Engineer

Nebius • United States

On-site
USD 120,000 - 160,000
Competitive compensation
Career growth
Flexibility & Ownership
+3