Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether

España

Presencial

EUR 90.000 - 150.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Remote-friendly Europe-based remote
GPU cloud exposure
NVIDIA GPUs

Descripción de la vacante

Jobgether, partnering with a leading AI infra provider in Spain, seeks a Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure). You will deploy and operate GPU-rich cloud environments, oversee datacenter readiness, and drive incidents to resolution across Linux, storage, networking, and virtualization.

The role combines hands-on engineering with cross-functional collaboration, requiring deep Linux, Kubernetes, and NCC/NVMe know-how, plus strong operational discipline in a

Formación

  • Datacenter infrastructure deployment experience.
  • Experience with GPU/AI infrastructure environments is a plus.
  • Strong Linux troubleshooting and system administration skills.
  • Networking knowledge including VLANs, BGP, and DC connectivity.
  • Automation and scripting (Terraform/Ansible/Python).
  • Observability and incident response experience.

Responsabilidades

  • Coordinate datacenter deployments with providers, integrators, and vendors.
  • Support rack-and-stack activities for GPU/CPU servers and networking gear.
  • Validate cabling, power, cooling, and physical readiness.
  • Commission GPU servers, storage nodes, and platform infra.
  • Perform burn-in, stress, and acceptance testing before production handover.
  • Validate RoCE/RDMA fabrics for AI workloads and troubleshoot issues.
  • Install and validate Ubuntu/Linux, NVIDIA drivers, CUDA, and container runtimes.
  • Support CloudStack, Kubernetes, KubeVirt, and CSI/CNI integrations.

Conocimientos

Datacenter infra
Linux troubleshooting
Networking
Kubernetes
Automation
Python
Observability
Communication

Herramientas

Terraform
Ansible
KVM/QEMU
Docker
Grafana
Prometheus

Descripción del empleo

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure) based in Spain.

This is ahighly hands-oninfrastructure role focusedon deploying, commissioning, and operating GPUcloud environments acrossregional and coredatacenters.

You will turnvalidated architecturesand bills ofmaterials into production-ready infrastructure spanninghardware, networking, storage, Linux, andplatform software.

The role sitsat the intersectionof datacenter operations, GPU infrastructure, networkengineering, and cloudplatform operations.

You will workwith high-densityNVIDIA GPU systems, advanced networking,storage platforms, Kubernetes, virtualization, and observabilitytooling.

As a practicaltechnical escalation point, you will troubleshootcomplex issues acrossphysical and softwarelayers and driveincidents through resolution.

You willalso help establishdeployment standards, validationprocedures, documentation, and operational practicesfor a rapidlyevolving AI infrastructureenvironment.

The role offersbroad technical ownershipinan international, fast-moving setting wherehands-on executionand operational excellenceare essential.

Accountabilities
  • Datacenter deployment:Coordinate deployments with datacenter providers, integrators, logistics teams, vendors, and internal engineering; validate rack layouts, power, cooling, airflow, cabling, labeling, and physical readiness.
  • Rack and infrastructure commissioning:Support rack-and-stack activities for GPU and CPU servers, storage, switches, routers, firewalls, PDUs, serial/OOB systems, and supporting infrastructure.
  • Cabling and connectivity:Validate fiber and copper cabling, optics, transceivers, breakout cables, port mappings, link speeds, redundancy, and management, storage, north-south, and east-west connectivity.
  • Hardware bring-up:Commission GPU servers, storage nodes, and platform infrastructure while validating BIOS, BMC, firmware, NICs, DPUs, GPUs, NVMe, RAID/HBA, PCIe topology, NUMA, thermals, power, and hardware health.
  • Hardware validation:Execute burn-in, stress, network, storage, and acceptance testing before production handover; troubleshoot issues involving GPUs, DPUs, NICs, optics, memory, disks, firmware, and BIOS.
  • Network deployment support:Work with network engineering to validate switch configurations, routing, VLAN/VRF segmentation, BGP, ECMP, EVPN/VXLAN, OVS/OVN, VyOS, firewalls, WAF infrastructure, and customer connectivity.
  • AI networking:Support validation of RoCE/RDMA fabrics for distributed AI workloads and troubleshoot issues such as link flaps, MTU mismatches, route errors, packet loss, PFC/ECN problems, and congestion.
  • Platform installation:Install and validate Ubuntu/Linux environments, NVIDIA drivers, CUDA, OFED or inbox drivers, Docker/containerd, KVM/QEMU, platform agents, and GPU infrastructure components.
  • Cloud and Kubernetes environments:Support CloudStack, Kubernetes, KubeVirt, GPU Operator, CSI/CNI integrations, GPU passthrough, SR-IOV, BlueField DPUs, VM networking, and container networking.
  • Storage integration:Support integration and validation of StorPool, Weka, local NVMe, and other supported storage platforms.
  • Operational readiness:Execute acceptance testing, produce deployment readiness reports, maintain runbooks, and ensure infrastructure is fully operational before customer or production handover.
  • Day-2 operations:Perform controlled firmware, OS, driver, BIOS, switch, and hardware maintenance while supporting production incidents and infrastructure escalations.
  • Incident management:Investigate operational failures, perform root-cause analysis, distinguish temporary workarounds from permanent fixes, and work with engineering to eliminate recurring issues.
  • Observability:Validate telemetry and monitoring across hosts, GPUs, DPUs, switches, storage, and platform components using tools such as Zabbix, Prometheus, Grafana, Loki, DCGM/NVML, and NVIDIA NetQ or equivalents.
  • Performance validation:Establish baselines for GPU, network, storage, and host performance and support benchmarking and infrastructure validation.
  • Documentation:Maintain accurate as-built records covering rack elevations, cable maps, port mappings, serial numbers, asset records, IP allocations, changes, and operational procedures.
  • Cross-functional coordination:Partner with infrastructure, networking, storage, platform, fleet automation, observability, product engineering, sales engineering, and service delivery teams.
  • Vendor management:Coordinate with datacenter providers, system integrators, server and storage vendors, NVIDIA, and networking suppliers to resolve deployment and infrastructure issues.
  • Continuous improvement:Feed field experience back into reference architectures, BOMs, rack designs, cabling standards, deployment playbooks, validation processes, and automation.
Requirements
  • Datacenter infrastructure:Strong hands-on experience deploying and maintaining datacenter infrastructure, ideally within GPU, HPC, AI cloud, private cloud, or high-density compute environments.
  • Bare-metal deployment:Proven ability to bring servers from physical installation and bare metal through validation and production readiness.
  • GPU infrastructure:Experience with NVIDIA GPU servers, drivers, firmware, PCIe topology, hardware validation, and high-performance compute environments.
  • Next-generation AI infrastructure:Familiarity with NVL72-style rack-scale architectures, NVLink/NVSwitch domains, in-rack networking, high-density power delivery, and OEM/NVIDIA validation requirements.
  • Datacenter readiness:Ability to assess power density, cooling, rack dimensions, floor loading, containment, serviceability, maintenance access, and other physical requirements for AI infrastructure.
  • Linux:Strong Linux troubleshooting capabilities and experience managing operating systems, kernels, drivers, and hardware interfaces.
  • Networking:Practical knowledge of VLANs, VRFs, BGP, ECMP, OVS/OVN, routing, OOB management, and high-speed datacenter connectivity.
  • GPU networking:Familiarity with NVIDIA/Mellanox networking, RoCE/RDMA, SR-IOV, BlueField DPUs, and high-performance east-west infrastructure.
  • Virtualization and containers:Experience with KVM/QEMU, VFIO, PCI passthrough, Docker/containerd, Kubernetes, and/or KubeVirt.
  • Storage:Experience integrating or troubleshooting local NVMe, storage nodes, and enterprise or distributed storage platforms.
  • Automation:Familiarity with Terraform, Ansible, Bash, and/or Python for deployment, validation, configuration, or operational automation.
  • Observability:Experience with infrastructure monitoring, telemetry, logs, metrics, health checks, and performance dashboards.
  • Documentation:Strong attention to detail and discipline in producing accurate as-built documentation, runbooks, validation records, and handover materials.
  • Troubleshooting:Strong systems-thinking ability across physical infrastructure, hardware, firmware, networking, Linux, storage, and platform layers.
  • Operational mindset:Comfortable supporting production environments, deployment windows, operational escalations, and customer-impacting incidents.
  • Communication:Able to clearly explain technical issues, risks, workarounds, and permanent solutions to engineering teams, vendors, and leadership.
  • Personal qualities:Highly practical, detail-oriented, calm under pressure, autonomous, and comfortable working both inside datacenters and remotely with smart-hands teams.
Benefits
  • Attractive compensation package reflecting your expertise, experience, transferable skills, and market conditions.
  • Full-time or contract engagement, depending on the agreed arrangement.
  • Europe-based remote working environment with flexibility.
  • Opportunity to work on cutting-edge GPU cloud and AI infrastructure at significant scale.
  • Hands-on exposure to NVIDIA GPU platforms, high-density datacenter environments, RoCE/RDMA networking, Kubernetes, virtualization, storage, and advanced observability.
  • Broad cross-functional scope spanning hardware, datacenter operations, networking, storage, Linux, and cloud platforms.
  • High-impact role within a fast-growing international scale-up.
  • Strong opportunities for technical growth and career development as the infrastructure platform expands.
  • Friendly, diverse, flexible, and international working environment.
  • Inclusive workplace committed to equal opportunity and respect for all qualified candidates.

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA • España

Presencial
EUR 90.000 - 130.000
Senior GPU Cloud Infra & Deployment Engineer
Senior GPU Cloud Infra & Deployment Engineer

Jobgether • España

Presencial
EUR 90.000 - 150.000
Remote-friendly Europe-based remote
GPU cloud exposure
NVIDIA GPUs
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)

Jobgether • España

Híbrido
EUR 120.000 - 170.000
Hybrid-friendly working environment
Career growth opportunities
International team
AI Infrastructure Engineer (GPU) - Remote EMEA
AI Infrastructure Engineer (GPU) - Remote EMEA

Pragmatike • Madrid

Presencial
EUR 60.000 - 80.000
Work from home flexibility
Inclusive recruitment process
Opportunity to influence core engineering decisions
Senior Product Designer - Compute, Nebius Console
Senior Product Designer - Compute, Nebius Console

Jobgether • España

Híbrido
EUR 70.000 - 100.000
Competitive compensation
Flexible work arrangements
Senior Network Engineer
Senior Network Engineer

Hamilton Barnes ? • España

Presencial
EUR 80.000 - 110.000
Data Center Delivery Manager (Nordics)
Data Center Delivery Manager (Nordics)

Volta • España

Presencial
EUR 155.000 - 215.000
Equity in Volta
Retirement plan
Health benefits
+1
Regional Key Account Manager – Neocloud
Regional Key Account Manager – Neocloud

Schneider Electric • Bellprat

Presencial
EUR 90.000 - 130.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Mirantis • Bellprat

Presencial
EUR 90.000 - 120.000
Competitive compensation package
Professional development and training
Conference attendance and tech talks
+1
Senior Developer Relations Manager - EMEA AI Natives
Senior Developer Relations Manager - EMEA AI Natives

NVIDIA • España

Presencial
EUR 90.000 - 215.000
Competitive salary
Generous benefits package
Career growth opportunities