A complete application in a minute — tailored resume and cover letter, ready to send.
NVIDIA DGX Cloud is seeking Senior Reliability Engineers to build automation, tooling, and operational systems for GPU clusters. You will work on Kubernetes-based infrastructure, Day 2 operability, and GitOps across DGX Cloud environments.
The role requires 8+ years in production infrastructure, strong Python/Go skills, expertise in Linux/Kubernetes, and a solid grasp of SRE principles. Phase 2 stands out with GPU-infra and observability focus.
NVIDIA DGX Cloud is seeking Senior Reliability Engineers to build automation, tooling, and operational systems for GPU clusters. You will work on Kubernetes-based infrastructure, Day 2 operability, and GitOps across DGX Cloud environments.
The role requires 8+ years in production infrastructure, strong Python/Go skills, expertise in Linux/Kubernetes, and a solid grasp of SRE principles. Phase 2 stands out with GPU-infra and observability focus.