Senior AI Cloud SRE: Kubernetes & GPU Work

NVIDIA AI

Zürich

Vor Ort

CHF 140.000 - 180.000

Vollzeit

vor 9 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

NVIDIA in Zürich is seeking a Senior Site Reliability Engineer to design, build, and operate scalable, reliable Kubernetes across major clouds for DGX Cloud. You will enhance monitoring, logging, and incident response for AI workloads and enterprise clients.

You will apply SRE principles, drive automation, and lead blameless postmortems while maintaining security and performance at scale. This role offers a challenging, high-impact environment in Switzerland.

Qualifikationen

  • BS in Computer Science or related field (or equivalent experience).
  • 10+ years of experience operating production services.
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
  • Experience with infrastructure automation tools (Terraform, Ansible, Chef, Puppet).
  • Proficiency in Python or Go.
  • In-depth knowledge of Linux, TCP/IP, and cloud security standards.
  • Proficient knowledge of SRE principles: SLOs, SLIs, error budgets, incident handling.
  • Experience building and operating observability stacks (OpenTelemetry, Prometheus, Grafana, ELK, Splunk).

Aufgaben

  • Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real time monitoring, logging and alerting.
  • Define SLOs/SLIs, monitor error budgets, and streamline reporting.
  • Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.
  • Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
  • Scale systems sustainably through automation and drive changes that improve reliability and velocity.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Practice balanced incident response and blameless postmortems.
  • Participate in on-call rotation to support production services

Kenntnisse

Kubernetes administration
Containerization
Microservices architecture
SRE principles
Observability stacks
Linux fundamentals
Networking (TCP/IP)
Programming (Python or Go)
Cloud security standards
Infrastructure automation tools

Ausbildung

BS in Computer Science or related field

Tools

Terraform
Ansible
Chef
Puppet

Jobbeschreibung

NVIDIA in Zürich is seeking a Senior Site Reliability Engineer to design, build, and operate scalable, reliable Kubernetes across major clouds for DGX Cloud. You will enhance monitoring, logging, and incident response for AI workloads and enterprise clients.

You will apply SRE principles, drive automation, and lead blameless postmortems while maintaining security and performance at scale. This role offers a challenging, high-impact environment in Switzerland.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior AI Cloud SRE | Kubernetes & Observability
Senior AI Cloud SRE | Kubernetes & Observability

NVIDIA • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA AI • Zürich

Vor Ort
CHF 140.000 - 180.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Gruppe • Zürich

Vor Ort
CHF 180.000 - 230.000
Senior HPC & AI Network Architect for Scalable AI Infra
Senior HPC & AI Network Architect for Scalable AI Infra

NVIDIA • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior HPC DevOps Engineer – Cloud, Automation, CI/CD
Senior HPC DevOps Engineer – Cloud, Automation, CI/CD

NVIDIA Gruppe • Zürich

Vor Ort
CHF 150.000 - 230.000
Senior HPC & AI Network Architect - Scale AI Infrastructure
Senior HPC & AI Network Architect - Scale AI Infrastructure

NVIDIA Switzerland AG • Zürich

Vor Ort
CHF 180.000 - 250.000
Competitive salary and benefits
Relocation assistance
Comprehensive benefits package
+1
Senior AI Network Architect for Scalable Distributed HPC
Senior AI Network Architect for Scalable Distributed HPC

NVIDIA Gruppe • Rüti (ZH)

Vor Ort
CHF 221.000 - 507.000
Comprehensive benefits package
Site Reliability Engineer: Scale Kubernetes & AI Ops
Site Reliability Engineer: Scale Kubernetes & AI Ops

Swissquote • Gland

Vor Ort
CHF 120.000 - 180.000
HPC-AI Cluster Architect for Next-Gen Systems
HPC-AI Cluster Architect for Next-Gen Systems

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000