Senior AI Cloud SRE | Kubernetes & Observability

NVIDIA

Zürich

Vor Ort

CHF 180.000 - 240.000

Vollzeit

Vor 12 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

NVIDIA in Zürich seeks an experienced Senior Site Reliability Engineer to maintain and scale DGX Cloud clusters for AI researchers and enterprise clients worldwide.

You will define SLOs, build observability stacks, automate infrastructure, and lead on-call incidents across multiple clouds, ensuring high reliability and velocity.

The role requires 10+ years of production experience, deep Kubernetes expertise, and strong programming skills in Python or Go.

Qualifikationen

  • BS in Computer Science or related technical field, or equivalent experience.
  • 10+ years of experience operating production services.
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices.
  • Experience with infrastructure automation tools (Terraform, Ansible, Chef, Puppet).
  • Proficiency in Python or Go.
  • Deep knowledge of Linux, TCP/IP, and cloud security.
  • Strong SRE fundamentals: SLOs/SLIs, error budgets, incident handling.
  • Experience with observability stacks: OpenTelemetry, Prometheus, Grafana, ELK, Splunk.

Aufgaben

  • Build, deploy and support large-scale Kubernetes clusters with emphasis on performance and reliability.
  • Define SLOs/SLIs and monitor error budgets, reporting outcomes.
  • Develop tools and automation for capacity planning and launch reviews.
  • Maintain live services by measuring availability, latency, and health.
  • Operate GPU workloads across cloud providers and private clouds.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Participate in on-call rotations and conduct blameless postmortems.

Kenntnisse

Kubernetes
SRE principles
Observability

Ausbildung

BS in Computer Science or related field

Tools

Terraform
Ansible
Chef
Puppet
OpenTelemetry
Prometheus
Grafana
ELK Stack
Lightstep
Splunk

Jobbeschreibung

NVIDIA in Zürich seeks an experienced Senior Site Reliability Engineer to maintain and scale DGX Cloud clusters for AI researchers and enterprise clients worldwide.

You will define SLOs, build observability stacks, automate infrastructure, and lead on-call incidents across multiple clouds, ensuring high reliability and velocity.

The role requires 10+ years of production experience, deep Kubernetes expertise, and strong programming skills in Python or Go.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior SRE: AI Cloud & Kubernetes Leader
Senior SRE: AI Cloud & Kubernetes Leader

NVIDIA Switzerland AG • Zürich

Vor Ort
CHF 150.000 - 230.000
Senior SRE: AI Cloud Infra & Kubernetes
Senior SRE: AI Cloud Infra & Kubernetes

NVIDIA Gruppe • Zürich

Vor Ort
CHF 180.000 - 230.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Switzerland AG • Zürich

Vor Ort
CHF 150.000 - 230.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Schweiz

Vor Ort
CHF 150.000 - 210.000
Senior SRE: AI-Driven Cloud Reliability Leader (Remote)
Senior SRE: AI-Driven Cloud Reliability Leader (Remote)

Jobgether • Schweiz

Vor Ort
CHF 43.000 - 96.000
100% remote work
Flexible working hours
Paid parental leave
Senior HPC AI Cluster Architect
Senior HPC AI Cluster Architect

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
HPC-AI Cluster Architect for Next-Gen Systems
HPC-AI Cluster Architect for Next-Gen Systems

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000
Senior Site Reliability Engineer - AI Infra & Multicloud
Senior Site Reliability Engineer - AI Infra & Multicloud

DeepJudge • Zürich

Vor Ort
CHF 120.000 - 180.000
Senior AI Network Architect for Scalable Distributed HPC
Senior AI Network Architect for Scalable Distributed HPC

NVIDIA Gruppe • Rüti (ZH)

Vor Ort
CHF 221.000 - 507.000
Comprehensive benefits package