Stand out for this role — generate a tailored resume and cover letter in about a minute.
HR POD Careers in Lahore seeks an experienced SRE to own and optimize production Linux GPU clusters and AI infrastructure. You will troubleshoot complex distributed systems, GPUs, network and storage issues, and steward high reliability across multi-node environments.
The role emphasizes automation, IaC with Ansible/Terraform, and modern tooling like Grafana/Prometheus. Expect shift-based hours, collaboration with Platform teams, and ongoing performance improvements.
8 PM - 4 AM