Senior SRE for GPU Clusters - Night Shift

HR POD Careers

Lahore

On-site

PKR 2,678,000 - 4,687,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

HR POD Careers is seeking a seasoned SRE to own the reliability and performance of production Linux GPU clusters in Lahore. You will lead complex multi-node troubleshooting, manage GPU drivers, networking, and storage, and automate provisioning and observability.

You will work with Kubernetes/Slurm, IaC tooling, and AI-assisted automation to reduce toil and improve SLAs; prior HPC experience and open-source contributions are a plus. This role operates on a night shift from 8 PM to 4 AM.

Qualifications

  • 5+ years of experience in systems, infrastructure, or SRE engineering.
  • Bachelor's degree in Computer Science.
  • Deep Linux troubleshooting across OS, networking, storage, and performance with root access on production systems.
  • Ubuntu experience highly relevant as primary OS.
  • Hands-on GPU server ops in production; driver and hardware issues.
  • Practical network troubleshooting incl. physical-layer faults.
  • Strong automation with Python or similar language.
  • Experience with configuration management and IaC (Ansible, Terraform etc.).
  • Observability/alerting with Grafana and Prometheus.
  • GPU cluster / AI infra production experience.
  • Kubernetes or Slurm in production; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with NVIDIA GPU stack, InfiniBand/RDMA, NCCL.
  • Experience with CLI-based AI coding agents like Claude Code.
  • Open-source contributions in cloud-native, HPC, or AI infra ecosystems.

Responsibilities

  • Own reliability, availability, and performance of production Linux GPU clusters (OS, drivers, GPUs, networking, storage).
  • Lead end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage.
  • Triage GPU NCCL and rail performance across multi-GPU/nodes.
  • Diagnose network faults end to end including cabling and link errors.
  • Configure/manage workload managers (Kubernetes, Slurm) and dependent services.
  • Build automation/tooling using Python/Go to reduce toil.
  • Leverage AI tools like Claude to accelerate automation and incident analysis.
  • Automate provisioning, image deployment, and remediation using Ansible/IaC.
  • Design observability with Grafana/Prometheus/Loki and tune meaningful alerts.
  • Lead incident response, on-call, postmortems, and reliability improvements.
  • Collaborate with Platform/Systems Eng for capacity planning and rollouts.

Skills

Deep Linux troubleshooting
Python
Go
Automation mindset
HPC/Research computing
NVIDIA GPU stack

Education

Bachelor's degree in Computer Science

Tools

Kubernetes
Slurm
Ansible
Terraform
Grafana
Prometheus
Loki
Claude Code

Job description

HR POD Careers is seeking a seasoned SRE to own the reliability and performance of production Linux GPU clusters in Lahore. You will lead complex multi-node troubleshooting, manage GPU drivers, networking, and storage, and automate provisioning and observability.

You will work with Kubernetes/Slurm, IaC tooling, and AI-assisted automation to reduce toil and improve SLAs; prior HPC experience and open-source contributions are a plus. This role operates on a night shift from 8 PM to 4 AM.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra SRE for Production Clusters
Senior GPU Infra SRE for Production Clusters

HR POD Careers • Lahore

On-site
PKR 3,000,000 - 4,200,000
GPU SRE — Production Linux & HPC Clusters (Night Shift)
GPU SRE — Production Linux & HPC Clusters (Night Shift)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 2,400,000 - 4,200,000
Site Reliability Engineer (Onsite, Lahore, PKR Salary)
Site Reliability Engineer (Onsite, Lahore, PKR Salary)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 2,400,000 - 4,200,000
Site Reliability Engineer (Onsite, Lahore, PKR Salary)
Site Reliability Engineer (Onsite, Lahore, PKR Salary)

HR POD Careers • Lahore

On-site
PKR 2,678,000 - 4,687,000
Site Reliability Engineer Onsite Lahore PKR Salary
Site Reliability Engineer Onsite Lahore PKR Salary

HR POD Careers • Lahore

On-site
PKR 3,000,000 - 4,200,000
Senior Linux Systems Engineer — Night Shift HPC & AI Infra
Senior Linux Systems Engineer — Night Shift HPC & AI Infra

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 725,000 - 949,000
Systems Engineer (Onsite, Lahore, PKR Salary)
Systems Engineer (Onsite, Lahore, PKR Salary)

HR POD Careers • Lahore

On-site
PKR 2,790,000 - 4,241,000
Senior Linux Systems Engineer – HPC & AI Infrastructure
Senior Linux Systems Engineer – HPC & AI Infrastructure

HR POD Careers • Lahore

On-site
PKR 2,790,000 - 4,241,000
Systems Engineer (Onsite, Lahore, PKR Salary)
Systems Engineer (Onsite, Lahore, PKR Salary)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 725,000 - 949,000
Systems Engineer Onsite Lahore PKR Salary
Systems Engineer Onsite Lahore PKR Salary

HR POD Careers • Lahore

On-site
PKR 360,000 - 600,000