Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking a Senior Site Reliability Engineer to join the BCM - DGX Cloud team. You will help deploy and operate large-scale GPU platforms, and bridge operations with development to improve reliability and performance.
You should have 8+ years in SRE/Software, strong Python, Linux, and networking skills. Experience with Slurm, Kubernetes, InfiniBand, and Spectrum-X is a plus. Base Command Manager expertise is valued.
NVIDIA is seeking a Senior Site Reliability Engineer to join the BCM - DGX Cloud team. You will help deploy and operate large-scale GPU platforms, and bridge operations with development to improve reliability and performance.
You should have 8+ years in SRE/Software, strong Python, Linux, and networking skills. Experience with Slurm, Kubernetes, InfiniBand, and Spectrum-X is a plus. Base Command Manager expertise is valued.