Site Reliability Engineer (Onsite, Lahore, PKR Salary)

HR POD Careers

Lahore

On-site

PKR 2,678,000 - 4,687,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

HR POD Careers is seeking a seasoned SRE to own the reliability and performance of production Linux GPU clusters in Lahore. You will lead complex multi-node troubleshooting, manage GPU drivers, networking, and storage, and automate provisioning and observability.

You will work with Kubernetes/Slurm, IaC tooling, and AI-assisted automation to reduce toil and improve SLAs; prior HPC experience and open-source contributions are a plus. This role operates on a night shift from 8 PM to 4 AM.

Qualifications

  • 5+ years of experience in systems, infrastructure, or SRE engineering.
  • Bachelor's degree in Computer Science.
  • Deep Linux troubleshooting across OS, networking, storage, and performance with root access on production systems.
  • Ubuntu experience highly relevant as primary OS.
  • Hands-on GPU server ops in production; driver and hardware issues.
  • Practical network troubleshooting incl. physical-layer faults.
  • Strong automation with Python or similar language.
  • Experience with configuration management and IaC (Ansible, Terraform etc.).
  • Observability/alerting with Grafana and Prometheus.
  • GPU cluster / AI infra production experience.
  • Kubernetes or Slurm in production; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with NVIDIA GPU stack, InfiniBand/RDMA, NCCL.
  • Experience with CLI-based AI coding agents like Claude Code.
  • Open-source contributions in cloud-native, HPC, or AI infra ecosystems.

Responsibilities

  • Own reliability, availability, and performance of production Linux GPU clusters (OS, drivers, GPUs, networking, storage).
  • Lead end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage.
  • Triage GPU NCCL and rail performance across multi-GPU/nodes.
  • Diagnose network faults end to end including cabling and link errors.
  • Configure/manage workload managers (Kubernetes, Slurm) and dependent services.
  • Build automation/tooling using Python/Go to reduce toil.
  • Leverage AI tools like Claude to accelerate automation and incident analysis.
  • Automate provisioning, image deployment, and remediation using Ansible/IaC.
  • Design observability with Grafana/Prometheus/Loki and tune meaningful alerts.
  • Lead incident response, on-call, postmortems, and reliability improvements.
  • Collaborate with Platform/Systems Eng for capacity planning and rollouts.

Skills

Deep Linux troubleshooting
Python
Go
Automation mindset
HPC/Research computing
NVIDIA GPU stack

Education

Bachelor's degree in Computer Science

Tools

Kubernetes
Slurm
Ansible
Terraform
Grafana
Prometheus
Loki
Claude Code

Job description

Requirements
  • 5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale.
  • Bachelor's degree in Computer Science.
  • Deep Linux troubleshooting skills across the OS, networking, storage, and performance, with hands-on experience working as root on production systems.
  • Experience with Ubuntu is highly relevant, as it is used almost exclusively.
  • Hands-on experience operating GPU servers in production, including troubleshooting driver, device, and hardware-level issues, rather than only the workloads running on top of them.
  • Practical network troubleshooting experience, including diagnosing physical-layer faults.
  • Strong automation mindset with programming skills in Python or a comparable language.
  • Experience with configuration management, node provisioning, and infrastructure-as-code (IaC) using Ansible, Terraform, or similar tools.
  • Experience building observability and alerting solutions using Grafana and Prometheus.
  • Experience operating GPU clusters or AI infrastructure at production scale.
  • Production experience with Kubernetes or Slurm; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with the NVIDIA GPU stack, InfiniBand/RDMA, and NCCL.
  • Experience with CLI-based AI coding agents such as Claude Code, rather than browser-based assistants alone.
  • Contributions to open-source projects within the cloud-native, HPC, or AI infrastructure ecosystem.
Responsibilities
  • Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage.
  • Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues.
  • Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
  • Diagnose network faults end to end, including configuration, routing, and physical-layer issues such as cabling, transceivers, and link errors.
  • Configure and maintain workload managers that schedule customer jobs, including Kubernetes, Slurm, or both, along with the identity, storage, and networking services they depend on.
  • Build automation and tooling to eliminate operational toil, using a modern language such as Python or Go to design solutions, review implementations, and redirect approaches when needed.
  • Use AI-assisted engineering tools such as Claude to accelerate automation, runbook development, and incident analysis.
  • Automate provisioning, image deployment, configuration, and remediation using Ansible and infrastructure-as-code.
  • Design and operate observability using Grafana, Prometheus, and Loki, while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
  • Lead incident response, on-call activities, blameless postmortems, and reliability improvements that maintain customer SLAs.
  • Partner with Platform and Systems Engineering teams on capacity planning, rollouts, and continuous improvement.
Working Hours

8 PM - 4 AM

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Onsite, Lahore, PKR Salary)
Site Reliability Engineer (Onsite, Lahore, PKR Salary)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 2,400,000 - 4,200,000
Systems Engineer (Onsite, Lahore, PKR Salary)
Systems Engineer (Onsite, Lahore, PKR Salary)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 725,000 - 949,000
Senior SRE for GPU Clusters - Night Shift
Senior SRE for GPU Clusters - Night Shift

HR POD Careers • Lahore

On-site
PKR 2,678,000 - 4,687,000
Systems Engineer
Systems Engineer

HR POD - Hiring Talent Globally • Lahore

On-site
PKR 2,500,000 - 4,000,000
Systems Engineer
Systems Engineer

Prime System Solutions • Lahore

On-site
PKR 2,500,000 - 4,200,000
Network Engineer (Onsite, Lahore, PKR Salary)
Network Engineer (Onsite, Lahore, PKR Salary)

HR POD Careers • Lahore

On-site
PKR 900,000 - 2,000,000
Network Engineer (Onsite, Lahore, PKR Salary)
Network Engineer (Onsite, Lahore, PKR Salary)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 2,790,000 - 3,906,000
GPU SRE — Production Linux & HPC Clusters (Night Shift)
GPU SRE — Production Linux & HPC Clusters (Night Shift)

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 2,400,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

Hopeghospital • Karachi Division

On-site
PKR 2,150,000 - 2,900,000
Senior Linux Systems Engineer — Night Shift HPC & AI Infra
Senior Linux Systems Engineer — Night Shift HPC & AI Infra

hr-pod-hiring-talent-globally • Lahore

On-site
PKR 725,000 - 949,000