Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI workloads. We seek a Senior Reliability Engineer to drive automation, tooling, and production systems across Kubernetes-based clusters and DGX environments.
You will develop observability stacks, define SLOs/SLIs, and partner with cross-functional teams to ensure production readiness. A strong background in SRE, distributed systems, and cloud-native ops is required.
NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI workloads. We seek a Senior Reliability Engineer to drive automation, tooling, and production systems across Kubernetes-based clusters and DGX environments.
You will develop observability stacks, define SLOs/SLIs, and partner with cross-functional teams to ensure production readiness. A strong background in SRE, distributed systems, and cloud-native ops is required.