Get more replies from employers
Send a job-specific resume in minutes.
hr-pod-hiring-talent-globally is seeking a senior SRE/engineer to own production Linux GPU clusters, drivers, and high-speed networking. You will troubleshoot complex distributed systems, multi-GPU paths, and NCCL performance, while managing workloads with Kubernetes or Slurm.
You will automate provisioning, image deployment, and remediation using Ansible and IaC, and build observability with Grafana/Prometheus. Collaboration with Platform teams drives reliability and capacity planning.
hr-pod-hiring-talent-globally is seeking a senior SRE/engineer to own production Linux GPU clusters, drivers, and high-speed networking. You will troubleshoot complex distributed systems, multi-GPU paths, and NCCL performance, while managing workloads with Kubernetes or Slurm.
You will automate provisioning, image deployment, and remediation using Ansible and IaC, and build observability with Grafana/Prometheus. Collaboration with Platform teams drives reliability and capacity planning.