Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA

España

Presencial

EUR 90.000 - 130.000

Jornada completa

Hace 8 días
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

NVIDIA is seeking a Senior Cloud Infrastructure and DevOps Solutions Architect to join its Infrastructure Specialist Team. You will engage with customers, partners, and cross‑functional teams to architect and guide large‑scale GPU‑accelerated infrastructure projects, spanning on‑prem and cloud, Kubernetes‑based platforms, and automation.

You will own day‑1 to day‑2 lifecycle, from hardware handover to production‑stable platforms, with emphasis on performance, reliability, and open‑source

Formación

  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
  • 8+ years in managing scalable cloud environments and automation engineering roles.
  • Cloud, HPC & GPU Expertise: Understanding networking fundamentals and data centre architectures, with hands-on experience managing HPC/AI clusters and NVIDIA GPU‑accelerated infrastructure—deployment, driver and CUDA toolkit management, optimisation, workload profiling and troubleshooting across CPUs, GPUs and high‑speed interconnects.
  • Kubernetes & AI/ML Workloads: Extensive background with Kubernetes, container orchestration, resource scheduling and scaling in GPU‑accelerated and HPC environments, including Slurm, KubeVirt, multi-tenant estates.
  • Linux & Storage Systems: Deep knowledge of Linux, OS security, and storage: Lustre, GPFS, ZFS, XFS, and Kubernetes storage tech.
  • Automation, GitOps & Observability: Proficiency in Python and Bash, IaC tools (Ansible, Terraform), GitOps lifecycle and upgrade management, and observability stacks (Grafana, Loki, Prometheus).
  • Fleet Reliability & Customer Engagement: Ability to measure MTBI and goodput on large GPU clusters, with consultative leadership and executive-facing communication.

Responsabilidades

  • Own full‑solution validation on the partner software stack, including cluster‑wide stability testing and multi‑day burn‑in against MTBI/goodput targets.
  • Minimise time from cluster handover to first production workload across hardware bring‑up, managed‑service intake and partner operations.
  • Own Day 2 production stability at fleet scale: monitoring, logging, orchestration, fault detection and remediation.
  • Assess customer environments and operate heterogeneous open platforms like Kubernetes, KubeVirt, Slurm, and GPU schedulers with enterprise networking/storage.
  • Provide consultative guidance and hands‑on troubleshooting across stack and support R&D, POCs and POVs validating new features and upgrade approaches.
  • Act as technical leader for assigned accounts: run knowledge transfer, create runbooks and onboarding materials for partner teams.

Conocimientos

Cloud infrastructure
DevOps
Kubernetes
Slurm
KubeVirt
Python
Bash scripting
Terraform
Ansible
GitOps
Observability
System reliability

Educación

BS/MS/PhD in CS or related fields

Herramientas

KubeVirt
Prometheus
Grafana
Lustre
GPFS
ZFS
XFS
Cumulus/SONiC
InfiniBand
NVLink/NVSwitch
DCGM

Descripción del empleo

NVIDIA is looking for a Senior Cloud Infrastructure and DevOps Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial organizations around the world are using NVIDIA products to redefine deep learning and data analytics, and to power next‑generation data centers. Join the team building and advising on many of the largest and fastest AI/HPC systems in the world!

We are looking for someone who combines deep technical expertise with strong consulting and communication skills. This role will engage directly with customers, partners, and cross‑functional teams to assess, architect, and guide the implementation of large‑scale infrastructure projects. The scope spans system architecture, Kubernetes‑based platforms, and automation—serving as both a trusted advisor and a hands‑on technical leader. You will sit at the centre of NVIDIA's Cloud Partner (NCP) operating model, covering the full Day 1 to Day 2 lifecycle: taking a GPU cluster from hardware handover, through full‑solution validation, to a production‑stable platform running at maximum goodput. NCP estates are open‑source‑first and heterogeneous—upstream Kubernetes, KubeVirt, Slurm, Prometheus/Grafana, Cumulus/SONiC and a long tail of ISV software—so this role is deliberately tool‑agnostic: you will meet each partner on the stack they actually run rather than on a single proprietary product.

What You’ll Be Doing
  • Own full‑solution validation on the partner software stack—the layer above hardware validation—including cluster‑wide stability testing, real training‑workload acceptance, and multi‑day, multi‑rack burn‑in against agreed MTBI and goodput targets.
  • Minimise the time from cluster handover to first production workload, working across hardware bring‑up, managed‑service intake and the partner's own operations teams to remove duplicated validation and handover friction.
  • Own Day 2 production stability at fleet scale: monitoring, logging and workload orchestration, fault detection and remediation, preventive maintenance, and proactive firmware and field‑notice rollout campaigns.
  • Assess customer environments and operate heterogeneous open platforms—upstream Kubernetes, KubeVirt, Slurm and GPU‑aware schedulers—integrated with enterprise‑grade networking and storage, and enable third‑party ISV workloads on top of them.
  • Provide consultative guidance and hands‑on troubleshooting across the full stack—bare metal, operating system, software stack, container platform, networking and storage—and support R&D, POCs and POVs validating new features, architectures and upgrade approaches.
  • Act as the technical leader for assigned accounts: run structured knowledge transfer and enablement, and produce runbooks, onboarding materials and best‑practice guides so partner teams can operate advanced configurations independently.
What We Need To See
  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
  • 8+ years in managing scalable cloud environments and automation engineering roles.
  • Cloud, HPC & GPU Expertise: Proven understanding of networking fundamentals and data centre architectures, with hands‑on experience managing HPC/AI clusters and NVIDIA GPU‑accelerated infrastructure—deployment, driver and CUDA toolkit management, optimisation, workload profiling and troubleshooting across CPUs, GPUs and high‑speed interconnects.
  • Kubernetes & AI/ML Workloads: Extensive background with Kubernetes for container orchestration, resource scheduling and scaling in GPU‑accelerated and HPC environments, including scheduler internals, batch schedulers such as Slurm, and mixed bare‑metal/virtualised (e.g. KubeVirt) multi‑tenant estates.
  • Linux & Storage Systems: Deep knowledge of Linux (RedHat, Ubuntu), OS‑level security, and protocols. Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and emerging Kubernetes storage technologies.
  • Automation, GitOps & Observability: Proficiency in Python and Bash scripting, configuration management and Infrastructure‑as‑Code tools (e.g. Ansible, Terraform), GitOps‑based cluster lifecycle and upgrade management for large fleets, and observability stacks (Grafana, Loki, Prometheus) for monitoring, logging and building fault‑tolerant systems.
  • Fleet Reliability & Customer Engagement: Demonstrated ability to measure and improve MTBI and job goodput on large GPU clusters—fault detection, drain and remediation workflows, SLO/error‑budget definition and post‑incident review—combined with a strong consultative background leading architectural reviews and presenting to executive stakeholders.
Ways To Stand Out From The Crowd
  • Knowledge of CI/CD pipelines and container‑based microservices architectures for software deployment and automation.
  • Experience with the NVIDIA GPU and Network Operators for automated GPU and network resource lifecycle management in Kubernetes, and with NVIDIA Base Command Manager (BCM) for provisioning, managing and monitoring GPU clusters at scale.
  • Familiarity with GPU health and fleet telemetry tooling—DCGM and XID diagnostics, node‑level health agents, and fleet‑wide reliability intelligence.
  • Expertise in AI‑native scheduling and inference frameworks on Kubernetes (e.g. KAI, Grove, Dynamo, NVIDIA Cloud Functions).
  • Background with RDMA‑based fabrics (InfiniBand or RoCE) in HPC or AI environments. Exposure to Cumulus Linux, SONiC or Spectrum‑X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations (NMX-C / NMX-M) on NVL72‑class systems is a strong plus.

, , JR2025837

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Engineer, NCX
Senior Engineer, NCX

NVIDIA • España

Presencial
EUR 110.000 - 150.000
Professional development opportunities
Senior Software and System Architect, Senior Software and System Architect
Senior Software and System Architect, Senior Software and System Architect

NVIDIA • Madrid

Presencial
EUR 60.000 - 90.000
Solution Architect, Local Government AI
Solution Architect, Local Government AI

NVIDIA • Madrid

Presencial
EUR 110.000 - 150.000
Data Center Engineer
Data Center Engineer

European Tech Recruit • España

Presencial
EUR 90.000 - 120.000
AI Infrastructure Solutions Engineer
AI Infrastructure Solutions Engineer

Ddn • Madrid

Híbrido
EUR 80.000 - 110.000
Network Engineer
Network Engineer

European Tech Recruit • España

Presencial
EUR 90.000 - 130.000
Director, Deployment Engineering — Systems Engineering
Director, Deployment Engineering — Systems Engineering

Nscale • Amer

Presencial
EUR 209.000 - 307.000
Senior HPC Cluster Engineer
Senior HPC Cluster Engineer

Nebius Group • Madrid

Presencial
EUR 80.000 - 120.000
Competitive compensation
Career growth
Flexible work
+3
Presales Workstation Technologist – Advanced Compute Solutions
Presales Workstation Technologist – Advanced Compute Solutions

Jobtailor • Barcelona

Presencial
EUR 60.000 - 90.000
Principal HPC Network Engineer (remote in the EU)
Principal HPC Network Engineer (remote in the EU)

Mirantis • Barcelona

Presencial
EUR 90.000 - 120.000