GPU SRE — Production Linux & HPC Clusters (Night Shift)

hr-pod-hiring-talent-globally

Lahore

On-site

PKR 2,400,000 - 4,200,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

hr-pod-hiring-talent-globally is seeking a senior SRE/engineer to own production Linux GPU clusters, drivers, and high-speed networking. You will troubleshoot complex distributed systems, multi-GPU paths, and NCCL performance, while managing workloads with Kubernetes or Slurm.

You will automate provisioning, image deployment, and remediation using Ansible and IaC, and build observability with Grafana/Prometheus. Collaboration with Platform teams drives reliability and capacity planning.

Qualifications

  • 5+ years in systems, infrastructure, or SRE managing production systems.
  • Deep Linux troubleshooting with Ubuntu experience.
  • Hands-on GPU server operations in production and driver issues.
  • Practical network troubleshooting incl. physical-layer faults.
  • Programming in Python or a similar language.
  • IaC experience with Ansible and Terraform.
  • Observability skills using Grafana and Prometheus.
  • Bachelor's degree in CS or an equivalent field.
  • Experience operating GPU clusters or AI infra in production.
  • Production experience with Kubernetes or Slurm; both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with NVIDIA GPU stack, InfiniBand/RDMA, NCCL.
  • Experience with Claude Code or AI coding agents.
  • Contributions to open-source projects in cloud-native/HPC/AI infra.

Responsibilities

  • Own reliability, availability, and performance of production Linux GPU clusters.
  • Lead end-to-end troubleshooting of distributed systems, GPU nodes, networking, and storage.
  • Troubleshoot GPU rail and NCCL performance across multi-GPU/multi-node paths.
  • Diagnose network faults end to end, including cabling and routing issues.
  • Configure and maintain workload managers (Kubernetes/Slurm) and identity/storage services.
  • Build automation using Python or Go to reduce toil and improve runbooks.
  • Use AI tools like Claude to accelerate automation and incident analysis.
  • Automate provisioning, image deployment, and configuration with IaC.
  • Design and operate observability with Grafana, Prometheus, Loki; tune alerts.
  • Lead on-call, postmortems, and reliability improvements to maintain SLAs.
  • Collaborate with Platform and Systems teams on capacity planning and rollouts.

Skills

GPU clusters
Linux administration
Kubernetes
Slurm
NVIDIA NCCL
IaC (Ansible Terraform)
Observability (Grafana Prometheus)
Python/Go
Networking

Education

Bachelor's degree in CS or related field

Tools

Ansible
Terraform
NVIDIA drivers
CUDA

Job description

hr-pod-hiring-talent-globally is seeking a senior SRE/engineer to own production Linux GPU clusters, drivers, and high-speed networking. You will troubleshoot complex distributed systems, multi-GPU paths, and NCCL performance, while managing workloads with Kubernetes or Slurm.

You will automate provisioning, image deployment, and remediation using Ansible and IaC, and build observability with Grafana/Prometheus. Collaboration with Platform teams drives reliability and capacity planning.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra SRE for Production Clusters
Senior GPU Infra SRE for Production Clusters

HR POD Careers • Lahore

On-site
PKR 3,000,000 - 4,200,000
Site Reliability Engineer (Onsite, Lahore, PKR Salary)
Site Reliability Engineer (Onsite, Lahore, PKR Salary)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 2,400,000 - 4,200,000
Site Reliability Engineer Onsite Lahore PKR Salary
Site Reliability Engineer Onsite Lahore PKR Salary

HR POD Careers • Lahore

On-site
PKR 3,000,000 - 4,200,000
Systems Engineer
Systems Engineer

Prime System Solutions • Lahore

On-site
PKR 2,500,000 - 4,200,000
Senior Linux Systems Engineer — Night Shift HPC & AI Infra
Senior Linux Systems Engineer — Night Shift HPC & AI Infra

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 725,000 - 949,000
Senior Linux Systems Engineer - AI/HPC Infra
Senior Linux Systems Engineer - AI/HPC Infra

HR POD Careers • Lahore

On-site
PKR 360,000 - 600,000
Senior Linux Systems Engineer - Scale and Automation
Senior Linux Systems Engineer - Scale and Automation

Prime System Solutions • Lahore

On-site
PKR 2,500,000 - 4,200,000
Senior Linux Systems Engineer – HPC & AI Infrastructure
Senior Linux Systems Engineer – HPC & AI Infrastructure

HR POD Careers • Lahore

On-site
PKR 2,790,000 - 4,241,000
Systems Engineer (Onsite, Lahore, PKR Salary)
Systems Engineer (Onsite, Lahore, PKR Salary)

HR POD Careers • Lahore

On-site
PKR 2,790,000 - 4,241,000
Staff/Senior DevOps Consultant - SRE (BI & Data Ecosystems)
Staff/Senior DevOps Consultant - SRE (BI & Data Ecosystems)

10Pearls, LLC • Lahore, Karachi Division

On-site
PKR 2,500,000 - 5,000,000