An application made for this job — a tailored resume and cover letter that speak straight to the posting.
NVIDIA DGX Cloud is seeking a Senior Reliability Engineer to design and operate automation for large GPU-based Kubernetes clusters across cloud partners and on-prem environments. You will shape observability, incident response, and production workflows to keep systems reliable and scalable.
You will implement SRE principles, define SLOs/SLIs, and collaborate across platform, storage, networking, and security teams. The role offers equity and a competitive total rewards package in Santa Clara, CA.
NVIDIA DGX Cloud is seeking a Senior Reliability Engineer to design and operate automation for large GPU-based Kubernetes clusters across cloud partners and on-prem environments. You will shape observability, incident response, and production workflows to keep systems reliable and scalable.
You will implement SRE principles, define SLOs/SLIs, and collaborate across platform, storage, networking, and security teams. The role offers equity and a competitive total rewards package in Santa Clara, CA.