An application made for this job — a tailored resume and cover letter that speak straight to the posting.
HR POD Careers is seeking a seasoned SRE to own the reliability and performance of production Linux GPU clusters in Lahore. You will lead complex multi-node troubleshooting, manage GPU drivers, networking, and storage, and automate provisioning and observability.
You will work with Kubernetes/Slurm, IaC tooling, and AI-assisted automation to reduce toil and improve SLAs; prior HPC experience and open-source contributions are a plus. This role operates on a night shift from 8 PM to 4 AM.
HR POD Careers is seeking a seasoned SRE to own the reliability and performance of production Linux GPU clusters in Lahore. You will lead complex multi-node troubleshooting, manage GPU drivers, networking, and storage, and automate provisioning and observability.
You will work with Kubernetes/Slurm, IaC tooling, and AI-assisted automation to reduce toil and improve SLAs; prior HPC experience and open-source contributions are a plus. This role operates on a night shift from 8 PM to 4 AM.