Remote GPU/HPC SRE for Large-Scale AI Inference

Saturn Cloud

United Arab Emirates

Remote

AED 360,000 - 600,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Saturn Cloud is seeking a GPU/HPC-focused Site Reliability Engineer for remote work across the MENA region. The role centers on GPU infrastructure, NVIDIA software, and Kubernetes-based workloads, with a focus on diagnosing and resolving production issues in large-scale GPU deployments.

You will work with NVIDIA drivers, NVML, and high-performance networking to ensure reliability and performance of inference workloads.

Qualifications

  • Hands-on GPU datacenter administration and driver troubleshooting.
  • Experience with Kubernetes GPU workloads and NVIDIA Operator.
  • Strong debugging skills for GPU, driver, and hardware issues.
  • Familiarity with CUDA, NVML, and nvidia-smi for diagnostics.

Responsibilities

  • Diagnose and resolve production issues on large-scale GPU inference infra.
  • Troubleshoot NVIDIA GPUs, drivers, CUDA compatibility, and GPU runtimes.
  • Investigate GPU health issues, Xid errors, and hardware/driver failures.
  • Analyze PCIe topology, NUMA, NVLink/NVSwitch, and GPU placement.
  • Troubleshoot multi-GPU and multi-node workloads in a Kubernetes context.
  • Work with GPU network fabric and distributed communication failures.
  • Differentiate between app, GPU, and infrastructure root causes; provide evidence.
  • Operate through Kubernetes-based GPU workloads and NVIDIA Operator plugins.

Skills

NVIDIA GPU administration
NVIDIA driver installation
CUDA compatibility
NVML and nvidia-smi
Kubernetes GPU Operator
Containerized GPU workloads
Prometheus/Grafana
Bash and Python
GPU health diagnostics
PCIe topology/NUMA
NVLink/NVSwitch
Networking (InfiniBand/RDMA)

Tools

NVIDIA DCGM
NVIDIA GPU Operator
Prometheus
Grafana
Kubernetes

Job description

Saturn Cloud is seeking a GPU/HPC-focused Site Reliability Engineer for remote work across the MENA region. The role centers on GPU infrastructure, NVIDIA software, and Kubernetes-based workloads, with a focus on diagnosing and resolving production issues in large-scale GPU deployments.

You will work with NVIDIA drivers, NVML, and high-performance networking to ensure reliability and performance of inference workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)
Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)

Saturn Cloud • United Arab Emirates

Remote
AED 360,000 - 600,000
Dubai AI Infra & Platform Engineer (Kubernetes + GPUs)
Dubai AI Infra & Platform Engineer (Kubernetes + GPUs)

Tech Junction Ltd • Dubai

On-site
AED 350,000 - 600,000
Technical Lead, GPU Infra (Remote)
Technical Lead, GPU Infra (Remote)

Tether Operations Limited • Abu Dhabi

Remote
AED 661,000 - 881,000
Remote Technical Lead, GPU Infrastructure Platform
Remote Technical Lead, GPU Infrastructure Platform

Tanqeeb • Dubai

On-site
AED 650,000 - 900,000
100% remote position.
Leadership and architecture ownership
Opportunity to work at scale in AI and
AI Infrastructure Engineer
AI Infrastructure Engineer

APPIT Software Inc. • Abu Dhabi

On-site
AED 180,000 - 240,000
GPU-HPC AI Infrastructure Engineer
GPU-HPC AI Infrastructure Engineer

APPIT Software Inc. • Abu Dhabi

On-site
AED 180,000 - 240,000
Senior MLOps Engineer - GPU Infra & Kubernetes
Senior MLOps Engineer - GPU Infra & Kubernetes

Sundus • Abu Dhabi

On-site
AED 300,000 - 600,000
DevOps / Infrastructure / SRE / Platform Engineering | Systemsltd | Dubai, Onsite
DevOps / Infrastructure / SRE / Platform Engineering | Systemsltd | Dubai, Onsite

Tech Junction Ltd • Dubai

On-site
AED 350,000 - 600,000
Senior AI Infrastructure Architect - GPU Compute & Kubernetes
Senior AI Infrastructure Architect - GPU Compute & Kubernetes

Open Innovation AI • Abu Dhabi

On-site
AED 420,000 - 640,000
Senior HPC Engineer – IFM
Senior HPC Engineer – IFM

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
AED 223,200 - 334,800