Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)

Saturn Cloud

United Arab Emirates

Remote

AED 360,000 - 600,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Saturn Cloud is seeking a GPU/HPC-focused Site Reliability Engineer for remote work across the MENA region. The role centers on GPU infrastructure, NVIDIA software, and Kubernetes-based workloads, with a focus on diagnosing and resolving production issues in large-scale GPU deployments.

You will work with NVIDIA drivers, NVML, and high-performance networking to ensure reliability and performance of inference workloads.

Qualifications

  • Hands-on GPU datacenter administration and driver troubleshooting.
  • Experience with Kubernetes GPU workloads and NVIDIA Operator.
  • Strong debugging skills for GPU, driver, and hardware issues.
  • Familiarity with CUDA, NVML, and nvidia-smi for diagnostics.

Responsibilities

  • Diagnose and resolve production issues on large-scale GPU inference infra.
  • Troubleshoot NVIDIA GPUs, drivers, CUDA compatibility, and GPU runtimes.
  • Investigate GPU health issues, Xid errors, and hardware/driver failures.
  • Analyze PCIe topology, NUMA, NVLink/NVSwitch, and GPU placement.
  • Troubleshoot multi-GPU and multi-node workloads in a Kubernetes context.
  • Work with GPU network fabric and distributed communication failures.
  • Differentiate between app, GPU, and infrastructure root causes; provide evidence.
  • Operate through Kubernetes-based GPU workloads and NVIDIA Operator plugins.

Skills

NVIDIA GPU administration
NVIDIA driver installation
CUDA compatibility
NVML and nvidia-smi
Kubernetes GPU Operator
Containerized GPU workloads
Prometheus/Grafana
Bash and Python
GPU health diagnostics
PCIe topology/NUMA
NVLink/NVSwitch
Networking (InfiniBand/RDMA)

Tools

NVIDIA DCGM
NVIDIA GPU Operator
Prometheus
Grafana
Kubernetes

Job description

This is a fully remote role open to candidates across the MENA region

About Saturn Cloud

Saturn Cloud builds infrastructure for running AI, machine learning, and data workloads at scale. Our platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure.

We’re looking for a GPU/HPC-focused Site Reliability Engineer based in the MENA region to help operate and troubleshoot the large-scale GPU infrastructure supporting Saturn Cloud Token Factory.

This role is focused on the infrastructure below and around the Kubernetes layer. You’ll serve as a technical escalation point for GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU performance issues affecting production inference workloads.

The ideal candidate comes from GPU infrastructure, HPC, AI infrastructure, neocloud, hyperscaler, or large-scale ML platform operations rather than traditional application support.

We value engineers who actively use modern AI coding agents to improve the speed and quality of their engineering and operational work.

What You’ll Do
  • Diagnose and resolve production issues affecting large-scale GPU inference infrastructure
  • Troubleshoot NVIDIA datacenter GPUs, drivers, CUDA compatibility, and GPU container runtimes
  • Investigate GPU health issues, including Xid errors and hardware or driver failure modes
  • Diagnose PCIe, NUMA, GPU placement, NVLink, and NVSwitch issues
  • Troubleshoot multi-GPU and multi-node workloads
  • Investigate high-performance networking and distributed communication failures
  • Distinguish application and inference-runtime issues from GPU, fabric, topology, node, driver, or hardware failures
  • Work with Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins
  • Use production observability and GPU metrics to diagnose reliability and performance issues
  • Work directly with GPU-cloud and infrastructure providers when incidents require hardware or fabric investigation
  • Produce clear technical evidence showing where the failure occurs and what the appropriate infrastructure team needs to investigate
What We’re Looking For
Core Skills
  • Hands-on experience using AI coding agents and agentic development tools to accelerate engineering, debugging, automation, and operational workflows
  • NVIDIA datacenter GPU administration
  • NVIDIA driver installation, upgrades, and troubleshooting
  • Strong understanding of CUDA and driver compatibility
  • Experience with NVML and nvidia-smi
  • GPU health diagnostics, including Xid errors and common hardware/driver failure modes
  • Understanding of PCIe topology, NUMA, and GPU placement
  • Understanding of NVLink and NVSwitch fundamentals
  • Familiarity with Kubernetes GPU Operator and device plugins
  • Experience running containerized GPU workloads
High-Performance Networking

You should have strong experience with several of the following:

  • InfiniBand
  • RDMA
  • RoCE
  • NCCL
  • GPUDirect RDMA
  • NCCL testing and distributed workload diagnostics
  • Bandwidth and latency troubleshooting

You should be able to determine whether a production issue originates in the application, GPU, network fabric, topology, node, or another underlying infrastructure layer.

Inference Infrastructure

You don’t need to be an ML researcher, but you should understand how modern inference workloads exercise GPU infrastructure.

  • vLLM, NVIDIA Dynamo, Triton, or comparable inference runtimes
  • Tensor and pipeline parallelism
  • Model loading and GPU memory consumptionKV cache
  • Continuous batching
  • GPU and memory utilization
  • Out-of-memory diagnosis
  • Multi-GPU and multi-node inference
  • Basic inference performance analysis, including throughput, latency, and GPU saturation
Additional Skills
  • Kubernetes troubleshooting sufficient to independently investigate GPU workloads inside a cluster
  • Prometheus/Grafana and NVIDIA DCGM metrics
  • Bash and Python
  • Experience with bare-metal GPU clusters or GPU cloud infrastructure
  • Familiarity with B200/B300/H200/H100-class systems is highly desirable
Strong Pluses
  • NVIDIA DCGM
  • NVIDIA Dynamo
  • KAI Scheduler or Grove
  • Spectrum-X
  • OFED/DOCA
  • Slurm or HPC cluster administration
  • Kubernetes-based GPU clouds
  • Large GPU fleet operations
  • GPU burn-in, qualification, and health-check tooling
What Success Looks Like

When an inference workload becomes unavailable or significantly slower, you can determine whether the problem is caused by the inference runtime, GPU memory pressure, a failed GPU, a driver/kernel interaction, PCIe/NVLink/NVSwitch topology, NCCL, RDMA/InfiniBand fabric, or the underlying node.

You can collect enough evidence to distinguish a Saturn Cloud software issue from an operator infrastructure or hardware problem and communicate precisely what an infrastructure provider needs to investigate.

You’re the person the team turns to when Kubernetes looks healthy, but the GPU workload isn’t.

Why Saturn Cloud
  • Remote-first culture with a high-trust, high-ownership environment
  • Work directly with cutting-edge GPU and AI infrastructure
  • Solve complex production problems spanning hardware, networking, Kubernetes, and modern inference systems
  • Help shape the reliability and operational practices behind large-scale AI workloads
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote GPU/HPC SRE for Large-Scale AI Inference
Remote GPU/HPC SRE for Large-Scale AI Inference

Saturn Cloud • United Arab Emirates

Remote
AED 360,000 - 600,000
Technical Lead - GPU Infrastructure (100% Remote - Worldwide) at Tether Operations Limited
Technical Lead - GPU Infrastructure (100% Remote - Worldwide) at Tether Operations Limited

Tether Operations Limited • Abu Dhabi

Remote
AED 661,000 - 881,000
DevOps / Infrastructure / SRE / Platform Engineering | Systemsltd | Dubai, Onsite
DevOps / Infrastructure / SRE / Platform Engineering | Systemsltd | Dubai, Onsite

Tech Junction Ltd • Dubai

On-site
AED 350,000 - 600,000
Senior HPC Engineer – IFM
Senior HPC Engineer – IFM

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
AED 223,200 - 334,800
Technical Lead, GPU Infra (Remote)
Technical Lead, GPU Infra (Remote)

Tether Operations Limited • Abu Dhabi

Remote
AED 661,000 - 881,000
Developer Technology Engineer, Energy
Developer Technology Engineer, Energy

NVIDIA • United Arab Emirates

On-site
AED 420,000 - 660,000
AI Infrastructure Engineer
AI Infrastructure Engineer

APPIT Software Inc. • Abu Dhabi

On-site
AED 180,000 - 240,000
Senior MLOps Engineer - GPU Infra & Kubernetes
Senior MLOps Engineer - GPU Infra & Kubernetes

Sundus • Abu Dhabi

On-site
AED 300,000 - 600,000
Agentic AI Engineer | Systems Ltd | Dubai, UAE
Agentic AI Engineer | Systems Ltd | Dubai, UAE

Systems Ltd • Dubai

On-site
AED 360,000 - 600,000
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • United Arab Emirates

Remote
AED 350,000 - 700,000
Remote-first
International team
Cutting-edge research
+2