An application made for this job — a tailored resume and cover letter that speak straight to the posting.
NVIDIA is seeking a Senior Systems Engineer for the Managed AI Research Superclusters (MARS) group to design and optimize batch scheduling for GPU clusters, improving resource fairness and performance across AI workloads.
You will work with SLURM/K8s schedulers, develop in C/C++, Go, Python, and container tech like Docker and Singularity, and help scale automation while solving complex reliability challenges.
NVIDIA is a pioneer in accelerated computing, known for inventing the GPU and driving breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence. Our technology powers everything from generative AI to autonomous systems, and we continue to shape the future of computing through innovation and collaboration.
Within this mission, our team, Managed AI Research Superclusters (MARS), builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next generation of AI/ML systems. By joining us, you'll help design solutions that power some of the world's most advanced computing workloads.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until August 17, 2026.
This posting is for an existing vacancy.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.