A complete application in a minute — tailored resume and cover letter, ready to send.
NVIDIA is seeking a Sr Site Reliability Engineer to help deploy and operate large-scale GPU platforms. You will bridge cluster operations with development, handle incidents, and design features in the Base Command Manager product. You will validate Slurm and Kubernetes configurations for performance, scale, and resilience.
Base salaries vary by level, with equity and benefits available. Applications accepted through August 27, 2026, for an existing vacancy.
NVIDIA is seeking a Sr Site Reliability Engineer to help deploy and operate large-scale GPU platforms. You will bridge cluster operations with development, handle incidents, and design features in the Base Command Manager product. You will validate Slurm and Kubernetes configurations for performance, scale, and resilience.
Base salaries vary by level, with equity and benefits available. Applications accepted through August 27, 2026, for an existing vacancy.