Stand out for this role — generate a tailored resume and cover letter in about a minute.
Nebius is seeking an experienced Site Reliability Engineer to own the reliability, performance, and observability of the full inference stack. You will design telemetry pipelines, tune Kubernetes autoscalers, and craft Terraform modules to ensure cost efficiency and resilience.
You’ll respond to incidents, drive post-morts, and collaborate with software engineers to turn reliability into a product feature. This role requires deep experience with Kubernetes, Prometheus, Grafana, Terraform, and
Nebius is seeking an experienced Site Reliability Engineer to own the reliability, performance, and observability of the full inference stack. You will design telemetry pipelines, tune Kubernetes autoscalers, and craft Terraform modules to ensure cost efficiency and resilience.
You’ll respond to incidents, drive post-morts, and collaborate with software engineers to turn reliability into a product feature. This role requires deep experience with Kubernetes, Prometheus, Grafana, Terraform, and