Senior Infrastructure Reliability Engineer: Day 2 Ops

NVIDIA

España

Presencial

EUR 110.000 - 150.000

Jornada completa

Hace 7 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Professional development opportunities

Descripción de la vacante

NVIDIA is seeking an NCX Senior Engineer to advance Day 2 operations for NVIDIA Cloud Partners within the DSX team. You’ll drive continuous validation, observability, and automated remediation across large AI clusters to ensure production reliability.

This hands-on role requires deep Kubernetes, Linux, and cloud experience, plus strong scripting in Python or Go and collaboration with partner teams to translate requirements into scalable operational practices.

Formación

  • BS, MS, or Ph.D. in CS/EE or related field, or equivalent experience.
  • 8+ years in infrastructure engineering, SRE, DevOps, or cloud production environments.
  • Strong Linux-based distributed systems and cloud infra production experience.
  • Deep Kubernetes, containers, cluster scheduling, and multi-node lifecycle knowledge.
  • Strong production observability experience (metrics, logs, alerts, dashboards, SLAs).
  • Automation for infra lifecycle, failure detection, remediation, upgrades, and config management.
  • Strong networking fundamentals and troubleshooting across compute, network, storage.
  • Programming and automation using Python, Go, or shell scripting.

Responsabilidades

  • Lead Day 2 readiness with NVIDIA Cloud Partners for systems, procedures, automation, and runbooks.
  • Build continuous infrastructure validation for GPU/CPU/storage/network in large AI clusters.
  • Establish observability and telemetry across compute, GPU, InfiniBand/RoCE, storage, Kubernetes, AI workloads.
  • Develop automated detection and remediation to return services after failures.
  • Refine fleet lifecycle administration: drivers, firmware, Kubernetes nodes, OS patches, config drift.
  • Operationalize NVIDIA reference architectures into production practices and runbooks.
  • Define health signals, SLOs, metrics, acceptance criteria for reliability and readiness.
  • Build reusable operational frameworks and reference implementations for multiple environments.

Conocimientos

Kubernetes
Linux
Observability
Automation
Python
Go
Shell scripting
Networking

Educación

BS/MS/PhD in Computer Science or Electrical Engineering or related field

Herramientas

Prometheus
Grafana
OpenTelemetry
Alertmanager

Descripción del empleo

NVIDIA is seeking an NCX Senior Engineer to advance Day 2 operations for NVIDIA Cloud Partners within the DSX team. You’ll drive continuous validation, observability, and automated remediation across large AI clusters to ensure production reliability.

This hands-on role requires deep Kubernetes, Linux, and cloud experience, plus strong scripting in Python or Go and collaboration with partner teams to translate requirements into scalable operational practices.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Engineer, NCX
Senior Engineer, NCX

NVIDIA • España

Presencial
EUR 110.000 - 150.000
Professional development opportunities
Senior Cloud Infra & DevOps Architect for AI/HPC
Senior Cloud Infra & DevOps Architect for AI/HPC

NVIDIA • España

Presencial
EUR 90.000 - 130.000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • España

Presencial
EUR 90.000 - 130.000
Director of Deployment & Platform Reliability
Director of Deployment & Platform Reliability

Nscale • Amer

Presencial
EUR 209.000 - 307.000
Remote Senior Platform Engineer - Multi-Cloud & AI Ops
Remote Senior Platform Engineer - Multi-Cloud & AI Ops

Allianz Commercial • Barcelona

Híbrido
EUR 90.000 - 130.000
Health insurance
Paid leave
Retirement plans
AI Infrastructure Solutions Engineer
AI Infrastructure Solutions Engineer

Ddn • Madrid

Híbrido
EUR 80.000 - 110.000
Senior SRE - AI Cloud Infra, Kubernetes & GPUs
Senior SRE - AI Cloud Infra, Kubernetes & GPUs

Mirantis • Bellprat

Presencial
EUR 90.000 - 120.000
Competitive compensation package
Professional development and training
Conference attendance and tech talks
+1
Senior Cloud SRE — Kubernetes, AWS & Reliability Lead
Senior Cloud SRE — Kubernetes, AWS & Reliability Lead

Nexthink • Madrid

Híbrido
EUR 60.000 - 86.000
Hybrid work model
Unlimited vacation
Volunteer days
+2
Senior SRE — Hybrid, Cloud-Native SaaS Reliability
Senior SRE — Hybrid, Cloud-Native SaaS Reliability

Nexthink • España

Híbrido
EUR 60.000 - 93.000
Hybrid work model
Unlimited vacation
Company-paid relocation
+2
Network Engineer
Network Engineer

European Tech Recruit • España

Presencial
EUR 90.000 - 130.000