Senior GPU Infra Engineer - Remote, Scale & Reliability

Sharon AI, Inc

United Arab Emirates

On-site

AED 420,000 - 660,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote-first across EMEA
Hands-on with GPU infrastructure
Impact on platform reliability

Job summary

Sharon AI, Inc. is seeking an experienced Senior Infrastructure Engineer to build, operate and improve a large-scale GPU‑based platform across multi‑tenant environments.

You will lead on Kubernetes, HPC schedulers, and Terraform/Ansible automation while partnering with ML and platform teams to optimize performance, reliability and cost.

This remote-first role spans across EMEA and offers the chance to shape production-grade infrastructure for AI workloads and future GPU services.

Qualifications

  • 6+ years in infrastructure engineering, platform engineering or SRE roles.
  • Deep expertise across cloud and/or bare-metal environments.
  • Hands-on with Kubernetes and container orchestration.
  • Experience with GPU-based systems and workloads.
  • Proficient in Terraform, Ansible, and similar IaC tools.
  • Programming/scripting in Python, Go, or Bash.

Responsibilities

  • Build and operate large-scale multi-tenant infrastructure for GPUaaS.
  • Manage HPC clusters and GPU-optimised infrastructure.
  • Develop Infrastructure-as-Code with Terraform/Ansible.
  • Automate provisioning, scaling and configuration.
  • Improve GPU utilisation, scheduling and cost efficiency.
  • Enhance observability with monitoring, logging, and alerts.
  • Ensure high availability, fault tolerance and DR capabilities.
  • Maintain security controls and workload isolation.
  • Collaborate with ML/platform teams to optimise performance.
  • Provide L3 support during standard working hours in EMEA.

Skills

Kubernetes
SRE
Cloud & Bare-metal
Python/Go/Bash
Security
Observability
Cost Optimization
Networking
Leadership
CI/CD
Distributed Systems

Education

Bachelor's degree in CS/Engineering

Tools

Terraform
Ansible
Slurm
GPU management tooling

Job description

Sharon AI, Inc. is seeking an experienced Senior Infrastructure Engineer to build, operate and improve a large-scale GPU‑based platform across multi‑tenant environments.

You will lead on Kubernetes, HPC schedulers, and Terraform/Ansible automation while partnering with ML and platform teams to optimize performance, reliability and cost.

This remote-first role spans across EMEA and offers the chance to shape production-grade infrastructure for AI workloads and future GPU services.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Prem AI Infra Engineer (Kubernetes & GPU)
On-Prem AI Infra Engineer (Kubernetes & GPU)

Evollabs Tech • Dubai

On-site
AED 480,000 - 720,000
AI Infrastructure Engineer
AI Infrastructure Engineer

APPIT Software Inc. • Abu Dhabi

On-site
AED 180,000 - 240,000
Senior HPC Engineer – IFM
Senior HPC Engineer – IFM

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
AI On-Prem Platform Engineer (Kubernetes & GPU)
AI On-Prem Platform Engineer (Kubernetes & GPU)

Evollabs Tech • Dubai Emirate

On-site
AED 300,000 - 520,000
GPU-HPC AI Infrastructure Engineer
GPU-HPC AI Infrastructure Engineer

APPIT Software Inc. • Abu Dhabi

On-site
AED 180,000 - 240,000
GPU Software Engineer - AI Acceleration
GPU Software Engineer - AI Acceleration

YO IT Consulting • Al Ruways Industrial City

On-site
AED 279,000 - 502,000
Remote contractor
Flexible hours
CUDA Engineering Expert - GPU Optimization
CUDA Engineering Expert - GPU Optimization

YO IT Consulting • Al Ruways Industrial City

On-site
AED 25,000 - 45,000
Senior AI-Native Infra & Reliability Engineer (Remote)
Senior AI-Native Infra & Reliability Engineer (Remote)

IgniteTech • United Arab Emirates

On-site
AED 250,000 - 400,000
Fully remote
Senior HPC Engineer: Build & Optimize AI Clusters
Senior HPC Engineer: Build & Optimize AI Clusters

Core42 • Abu Dhabi

On-site
AED 420,000 - 650,000
Competitive Salary
Yearly Bonus
Discount Cards Esaad and Fazaa
+2
Remote Performance Engineer for AI Systems Optimization
Remote Performance Engineer for AI Systems Optimization

Hire Feed • Dubai

On-site
AED 354,000 - 557,000