Global Network SRE for AI Cloud Infrastructure

Socket.dev

United States

On-site

USD 180,000 - 224,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth
Flexibility and ownership
Collaborative culture
Impactful AI projects
International teams

Job summary

Nebius is hiring a Network Site Reliability Engineer to define reliability targets, build automation, and run the Network infrastructure at scale. This engineering-first SRE role focuses on creating safe, observable, and efficient network operations to support a rapidly growing AI cloud platform.

You will own incidents, drive improvements, and partner with platform teams to embed operability into design and deployment, enabling fast and safe changes across global networks.

Qualifications

  • Strong production Linux fundamentals and a structured approach to debugging complex systems.
  • Solid understanding of networking basics and how real networks fail.
  • Hands-on experience operating high-availability systems and improving them over time.
  • Ability to write and maintain software/automation (Go is common; Python is welcome).
  • Experience with modern infrastructure tooling (IaC, CI/CD, container platforms).

Responsibilities

  • Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets).
  • Drive reliability improvements across the network, including site readiness and inter-site connectivity.
  • Own incident response, lead investigations, and implement durable fixes.
  • Build and evolve observability with metrics/logs/traces, alerts, and rapid debugging.
  • Design safer change workflows with automation, CI/CD, test environments, and canaries.
  • Collaborate with network engineers and platform teams to embed operability in designs.

Skills

Linux fundamentals
Networking basics
High-availability systems
Automation (Go/Python)
IaC / CI-CD / containers

Job description

Nebius is hiring a Network Site Reliability Engineer to define reliability targets, build automation, and run the Network infrastructure at scale. This engineering-first SRE role focuses on creating safe, observable, and efficient network operations to support a rapidly growing AI cloud platform.

You will own incidents, drive improvements, and partner with platform teams to embed operability into design and deployment, enabling fast and safe changes across global networks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

Socket.dev • United States

On-site
USD 180,000 - 224,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Remote SRE — AI Cloud Hardware Infra
Remote SRE — AI Cloud Hardware Infra

Nebius • United States

On-site
USD 130,000 - 180,000
Health insurance
401(k) plan
Parental leave
+2
Remote SRE, Hardware Infra for AI Cloud
Remote SRE, Hardware Infra for AI Cloud

Nebius • United States

On-site
USD 130,000 - 180,000
100% company-paid medical, dental, and vision insurance
401(k) plan with company match
20 weeks paid parental leave for primary caregivers
+2
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior AI Cloud Network Engineer
Senior AI Cloud Network Engineer

Nebius • United States

On-site
USD 125,000 - 180,000
Comprehensive health insurance
401(k) plan with company contribution
Paid time off and public holidays
+2
Senior Data Center Network Architect for AI Cloud
Senior Data Center Network Architect for AI Cloud

Nebius • United States

On-site
USD 125,000 - 180,000
Health insurance
401(k) plan with company contribution
Paid time off
Senior Network Reliability Engineer | Platform SRE
Senior Network Reliability Engineer | Platform SRE

Group 1001 • Indianapolis (IN)

On-site
USD 135,000 - 190,000
Health insurance
Dental insurance
Vision insurance
+4
Staff Software Engineer, AI Cloud Infrastructure
Staff Software Engineer, AI Cloud Infrastructure

Nebius • United States

On-site
USD 175,000 - 225,000
Health insurance
401(k) plan
Parental leave
+2
Cloud Network SRE: Secure, Scalable Networking & IaC
Cloud Network SRE: Secure, Scalable Networking & IaC

Cerebras • San Francisco (CA)

On-site
USD 157,000 - 239,000
Excellent medical/dental/vision plans
401(k) plan and equity options
Relocation assistance
+5
Remote Senior Network Reliability Engineer (SRE)
Remote Senior Network Reliability Engineer (SRE)

Gainbridge • Zionsville (IN), Northern (KY)

On-site
USD 135,000 - 190,000
Health Insurance
Dental Insurance
Vision Insurance
+4