Senior Cloud Infra & DevOps Architect for AI/HPC

NVIDIA

España

Presencial

EUR 90.000 - 130.000

Jornada completa

Hace 8 días
Generador de candidaturas

Una candidatura completa en un minuto — currículum adaptado y carta de presentación, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

NVIDIA is seeking a Senior Cloud Infrastructure and DevOps Solutions Architect to join its Infrastructure Specialist Team. You will engage with customers, partners, and cross‑functional teams to architect and guide large‑scale GPU‑accelerated infrastructure projects, spanning on‑prem and cloud, Kubernetes‑based platforms, and automation.

You will own day‑1 to day‑2 lifecycle, from hardware handover to production‑stable platforms, with emphasis on performance, reliability, and open‑source

Formación

  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
  • 8+ years in managing scalable cloud environments and automation engineering roles.
  • Cloud, HPC & GPU Expertise: Understanding networking fundamentals and data centre architectures, with hands-on experience managing HPC/AI clusters and NVIDIA GPU‑accelerated infrastructure—deployment, driver and CUDA toolkit management, optimisation, workload profiling and troubleshooting across CPUs, GPUs and high‑speed interconnects.
  • Kubernetes & AI/ML Workloads: Extensive background with Kubernetes, container orchestration, resource scheduling and scaling in GPU‑accelerated and HPC environments, including Slurm, KubeVirt, multi-tenant estates.
  • Linux & Storage Systems: Deep knowledge of Linux, OS security, and storage: Lustre, GPFS, ZFS, XFS, and Kubernetes storage tech.
  • Automation, GitOps & Observability: Proficiency in Python and Bash, IaC tools (Ansible, Terraform), GitOps lifecycle and upgrade management, and observability stacks (Grafana, Loki, Prometheus).
  • Fleet Reliability & Customer Engagement: Ability to measure MTBI and goodput on large GPU clusters, with consultative leadership and executive-facing communication.

Responsabilidades

  • Own full‑solution validation on the partner software stack, including cluster‑wide stability testing and multi‑day burn‑in against MTBI/goodput targets.
  • Minimise time from cluster handover to first production workload across hardware bring‑up, managed‑service intake and partner operations.
  • Own Day 2 production stability at fleet scale: monitoring, logging, orchestration, fault detection and remediation.
  • Assess customer environments and operate heterogeneous open platforms like Kubernetes, KubeVirt, Slurm, and GPU schedulers with enterprise networking/storage.
  • Provide consultative guidance and hands‑on troubleshooting across stack and support R&D, POCs and POVs validating new features and upgrade approaches.
  • Act as technical leader for assigned accounts: run knowledge transfer, create runbooks and onboarding materials for partner teams.

Conocimientos

Cloud infrastructure
DevOps
Kubernetes
Slurm
KubeVirt
Python
Bash scripting
Terraform
Ansible
GitOps
Observability
System reliability

Educación

BS/MS/PhD in CS or related fields

Herramientas

KubeVirt
Prometheus
Grafana
Lustre
GPFS
ZFS
XFS
Cumulus/SONiC
InfiniBand
NVLink/NVSwitch
DCGM

Descripción del empleo

NVIDIA is seeking a Senior Cloud Infrastructure and DevOps Solutions Architect to join its Infrastructure Specialist Team. You will engage with customers, partners, and cross‑functional teams to architect and guide large‑scale GPU‑accelerated infrastructure projects, spanning on‑prem and cloud, Kubernetes‑based platforms, and automation.

You will own day‑1 to day‑2 lifecycle, from hardware handover to production‑stable platforms, with emphasis on performance, reliability, and open‑source

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • España

Presencial
EUR 90.000 - 130.000
Senior Engineer, NCX
Senior Engineer, NCX

NVIDIA • España

Presencial
EUR 110.000 - 150.000
Professional development opportunities
Senior Software and System Architect, Senior Software and System Architect
Senior Software and System Architect, Senior Software and System Architect

NVIDIA • Madrid

Presencial
EUR 60.000 - 90.000
Remote AI Network Architect for GPU Data Centers
Remote AI Network Architect for GPU Data Centers

Hamilton Barnes ? • España

Presencial
EUR 120.000 - 180.000
AI Infrastructure Solutions Architect - Pre‑Sales
AI Infrastructure Solutions Architect - Pre‑Sales

Schneider Electric • Barcelona

Presencial
EUR 90.000 - 130.000
Solution Architect, Local Government AI
Solution Architect, Local Government AI

NVIDIA • Madrid

Presencial
EUR 110.000 - 150.000
Senior HPC GPU Cluster Engineer
Senior HPC GPU Cluster Engineer

Nebius Group • Madrid

Presencial
EUR 80.000 - 120.000
Competitive compensation
Career growth
Flexible work
+3
Remote Technical Lead - GPU Infrastructure Platform
Remote Technical Lead - GPU Infrastructure Platform

Tether • Barcelona

Presencial
EUR 90.000 - 130.000
Senior Data Center Technician
Senior Data Center Technician

WeEngage Group | B Corp™ • Madrid

Presencial
EUR 55.000 - 75.000
AI Infrastructure Solutions Engineer
AI Infrastructure Solutions Engineer

Ddn • Madrid

Híbrido
EUR 80.000 - 110.000