Network SRE: Reliability & Automation for Cloud Infra

Nebius

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nebius is seeking a Site Reliability Engineer to build and run the Network infrastructure for a global AI cloud platform. You will set reliability targets, create tooling, and automate operations to scale safely.

Responsibilities include defining SLIs/SLOs, driving improvements across inter-site connectivity, owning incident responses, and evolving observability. Strong Linux, networking, and automation skills are required.

Qualifications

  • Strong production Linux fundamentals and a structured approach to debugging complex systems.
  • Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency, loss).
  • Hands-on experience operating high-availability systems and improving them over time.
  • Ability to write and maintain software/automation (Go is common; Python is also welcome).
  • Experience with modern infrastructure tooling (IaC, CI/CD, container platforms) and automating operations workflows.

Responsibilities

  • Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets).
  • Drive reliability improvements across the network including site readiness and inter-site connectivity.
  • Own incident response, lead investigations and postmortems, implement durable fixes.
  • Build and evolve observability with metrics/logs/traces, alerts, and faster debugging loops.
  • Design safer change workflows with automation, CI/CD, test/staging environments, canaries and rollbacks.

Skills

Linux fundamentals
Networking fundamentals
HA systems experience
Go
Python
IaC
CI/CD
Container platforms
Observability tooling

Tools

CI/CD tools
IaC tooling
Container platforms

Job description

Nebius is seeking a Site Reliability Engineer to build and run the Network infrastructure for a global AI cloud platform. You will set reliability targets, create tooling, and automate operations to scale safely.

Responsibilities include defining SLIs/SLOs, driving improvements across inter-site connectivity, owning incident responses, and evolving observability. Strong Linux, networking, and automation skills are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE, Hardware Infra for AI Cloud
Remote SRE, Hardware Infra for AI Cloud

Nebius • United States

On-site
USD 130,000 - 180,000
100% company-paid medical, dental, and vision insurance
401(k) plan with company match
20 weeks paid parental leave for primary caregivers
+2
Remote SRE — AI Cloud Hardware Infra
Remote SRE — AI Cloud Hardware Infra

Nebius • United States

On-site
USD 130,000 - 180,000
Health insurance
401(k) plan
Parental leave
+2
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops
Senior SRE: Cloud Reliability, CI/CD & High-Load Ops

Nebius • United States

Remote
USD 120,000 - 170,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior SRE - Compute Nodes (Linux & Virtualization)
Senior SRE - Compute Nodes (Linux & Virtualization)

Nebius • United States

Remote
USD 140,000 - 190,000
Remote Senior Network Reliability Engineer (SRE)
Remote Senior Network Reliability Engineer (SRE)

Gainbridge • Zionsville (IN), Northern (KY)

On-site
USD 135,000 - 190,000
Health Insurance
Dental Insurance
Vision Insurance
+4
Cloud SRE: Build Resilient, Secure, Scalable Infra
Cloud SRE: Build Resilient, Secure, Scalable Infra

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Cloud Network SRE: Secure, Scalable Networking & IaC
Cloud Network SRE: Secure, Scalable Networking & IaC

Cerebras • San Francisco (CA)

On-site
USD 157,000 - 239,000
Excellent medical/dental/vision plans
401(k) plan and equity options
Relocation assistance
+5
Senior Network Reliability Engineer | Platform SRE
Senior Network Reliability Engineer | Platform SRE

Group 1001 • Indianapolis (IN)

On-site
USD 135,000 - 190,000
Health insurance
Dental insurance
Vision insurance
+4
Senior Network Reliability Engineer, SRE for Cloud Platform
Senior Network Reliability Engineer, SRE for Cloud Platform

Gainbridge • United States

Hybrid
USD 135,000 - 190,000
Health insurance
Dental insurance
401(k) plan with company matching
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000