Senior GPU Cloud SRE - Scale & Observability

NVIDIA Corporation

Zürich

Vor Ort

CHF 150.000 - 190.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Verschicke keinen 08/15-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

NVIDIA Corporation is hiring a Senior Site Reliability Engineer to manage large-scale Kubernetes clusters, optimize GPU workloads, and uphold reliability across clouds. You will define SLOs/SLIs, lead incident response, and build tools to monitor latency, health and performance for researchers and enterprise clients.

Ideal candidates have 10+ years in production operations, strong Linux networking, cloud security, and hands-on experience with Terraform, Ansible, and observability stacks.

Qualifikationen

  • BS in Computer Science or related field, or equivalent experience.
  • 10+ years operating production services.
  • Expert Kubernetes admin, containerization and microservices experience.
  • Proficiency with Terraform/Ansible/Chef/Puppet for infra automation.

Aufgaben

  • Build, implement and support reliability aspects of large-scale Kubernetes clusters.
  • Define SLOs/SLIs, monitor error budgets, and streamline reporting.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Operate GPU workloads across multiple clouds and private clouds.
  • Maintain observability stacks using OpenTelemetry, Prometheus, Grafana, ELK, Splunk.

Kenntnisse

Kubernetes
Containerization
Microservices
Python
Go
Linux networking
Cloud security
SRE principles
Observability
Incident handling

Ausbildung

BS in Computer Science

Tools

Terraform
Ansible
Chef
Puppet
OpenTelemetry
Prometheus
Grafana
ELK/Elastic
Lightstep
Splunk

Jobbeschreibung

NVIDIA Corporation is hiring a Senior Site Reliability Engineer to manage large-scale Kubernetes clusters, optimize GPU workloads, and uphold reliability across clouds. You will define SLOs/SLIs, lead incident response, and build tools to monitor latency, health and performance for researchers and enterprise clients.

Ideal candidates have 10+ years in production operations, strong Linux networking, cloud security, and hands-on experience with Terraform, Ansible, and observability stacks.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Schweiz

Vor Ort
CHF 150.000 - 210.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Gruppe • Zürich

Vor Ort
CHF 180.000 - 230.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Corporation • Zürich

Vor Ort
CHF 150.000 - 190.000
Remote Senior Performance Engineer - AI & HPC Systems
Remote Senior Performance Engineer - AI & HPC Systems

NVIDIA Corporation • Zürich

Vor Ort
CHF 140.000 - 210.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA Gruppe • Zürich

Vor Ort
CHF 120.000 - 180.000
Senior HPC-AI Cluster Architect
Senior HPC-AI Cluster Architect

NVIDIA Gruppe • Zürich

Vor Ort
CHF 120.000 - 180.000
Site Reliability Engineer: Scale Kubernetes & AI Ops
Site Reliability Engineer: Scale Kubernetes & AI Ops

Swissquote • Gland

Vor Ort
CHF 120.000 - 180.000
Senior SRE - Hybrid Cloud for Real-Time Sports
Senior SRE - Hybrid Cloud for Real-Time Sports

Geniussports • Lausanne

Hybrid
CHF 140.000 - 190.000
Senior HPC AI Network Architect: Scalable Infra
Senior HPC AI Network Architect: Scalable Infra

NVIDIA Corporation • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior GPU Networking Architect for AI Systems
Senior GPU Networking Architect for AI Systems

NVIDIA Corporation • Zürich

Vor Ort
CHF 180.000 - 240.000
Competitive salary
Comprehensive benefits package