Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Nebius is seeking a Site Reliability Engineer to own the Network infrastructure, defining reliability targets and building tooling to scale reliably. This engineering-first role focuses on measurable SLIs/SLOs, robust change workflows, and strong operability across the global network.
You will collaborate with network engineers and platform teams to embed observability, automate incident response, and ensure rapid recovery from issues as Nebius expands in North America and beyond.
About Nebius : Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
We’re looking for a Site Reliability Engineer to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly. Your responsibilities will include:
We expect you to have: Strong p