Senior SRE: AI Cloud Infra & Kubernetes

NVIDIA Gruppe

Zürich

Presencial

CHF 180.000 - 230.000

Jornada completa

14 días+
Generador de candidaturas

Destaca para este puesto — genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Descripción de la vacante

NVIDIA is seeking a Senior Site Reliability Engineer to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide. You will build and support Kubernetes-based systems, define SLOs/SLIs, and drive reliability across multi-cloud environments.

Ideal candidates have 10+ years operating production services, deep Kubernetes and observability expertise, and hands-on automation experience with Terraform/Ansible.

Formación

  • BS in Computer Science or a related technical field or equivalent experience.
  • 10+ years of experience operating production services.
  • Expert knowledge of Kubernetes administration, containerization, and microservices.
  • Experience with infrastructure automation tools (Terraform, Ansible, Chef, Puppet).
  • Proficiency in Python or Go.
  • Deep Linux networking, cloud security fundamentals.
  • Strong knowledge of SRE principles: SLOs/SLIs, error budgets, incident handling.
  • Experience building observability stacks (Prometheus, Grafana, OpenTelemetry, ELK).

Responsabilidades

  • Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real time monitoring, logging and alerting.
  • Define SLOs/SLIs, monitor error budgets, and streamline reporting.
  • Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.
  • Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
  • Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Practice balanced incident response and blameless postmortems.
  • Participate in on-call rotation to support production services.

Conocimientos

Kubernetes administration
Containerization
Microservices architecture
Terraform
Ansible
Python
Go
Linux networking
Cloud security
SRE principles
Observability stacks

Educación

BS in Computer Science or related field

Herramientas

Terraform
Ansible
Chef
Puppet
OpenTelemetry
Prometheus
Grafana
ELK Stack

Descripción del empleo

NVIDIA is seeking a Senior Site Reliability Engineer to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide. You will build and support Kubernetes-based systems, define SLOs/SLIs, and drive reliability across multi-cloud environments.

Ideal candidates have 10+ years operating production services, deep Kubernetes and observability expertise, and hands-on automation experience with Terraform/Ansible.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Suiza

Presencial
CHF 150.000 - 210.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Corporation • Zürich

Presencial
CHF 180.000 - 230.000
Senior Production Engineer - AI Cloud Reliability
Senior Production Engineer - AI Cloud Reliability

NVIDIA • Zürich

Presencial
CHF 140.000 - 180.000
Senior Production Engineer - DGX Cloud
Senior Production Engineer - DGX Cloud

NVIDIA • Zürich

Presencial
CHF 140.000 - 180.000
Senior SRE Engineer, Prod AI & Cloud-Scale Systems
Senior SRE Engineer, Prod AI & Cloud-Scale Systems

Google LLC • Zürich

Presencial
CHF 120.000 - 170.000
Site Reliability Engineer: Scale Kubernetes & AI Ops
Site Reliability Engineer: Scale Kubernetes & AI Ops

Swissquote • Gland

Presencial
CHF 120.000 - 180.000
Remote Senior Performance Engineer - AI & HPC Systems
Remote Senior Performance Engineer - AI & HPC Systems

NVIDIA Corporation • Zürich

Presencial
CHF 140.000 - 210.000
Senior Site Reliability Engineer - AI Infra & Multicloud
Senior Site Reliability Engineer - AI Infra & Multicloud

DeepJudge • Zürich

Presencial
CHF 120.000 - 180.000
Senior HPC AI Network Architect: Scalable Infra
Senior HPC AI Network Architect: Scalable Infra

NVIDIA Corporation • Zürich

Presencial
CHF 180.000 - 240.000
Site Reliability Engineer — AI Search Infra on Multi-Cloud
Site Reliability Engineer — AI Search Infra on Multi-Cloud

DeepJudge AG • Zürich

Presencial
CHF 120.000 - 180.000