Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA in Santa Clara, CA is seeking a System Software Engineer, Performance for the CUDA driver and runtime.
You will design and ship production C/C++ features, optimize key execution paths, and trace workloads across the CPU, OS, and GPU boundaries to improve latency and throughput.
Collaborate with multiple teams to set direction, bring up new platforms, and influence future CUDA architectures while growing your expertise in deep systems software.
The AI revolution is not powered by models alone, rather it advances when enormous amounts of computation become fast, efficient, and economical enough to turn new ideas into products people can use on a global scale. Faster training lets research and product teams test the next idea sooner. Lower-latency, higher-throughput inference makes AI assistants and agents more responsive and practical for more people. Shorter time to solution lets scientists and engineers explore more possibilities within the same time and energy budget.
At NVIDIA, performance is not a supporting metric - it is how architectural invention becomes useful computing. CUDA is a critical layer where that transformation happens, sitting beneath the frameworks, libraries, and applications used across AI, deep learning, and HPC, as well as graphics, automotive, robotics, and other CUDA-powered products. That gives this team unusual leverage: reduce overhead in a fundamental launch, synchronization, memory, or data-movement path-or create a new driver or runtime capability-and the improvement can flow through many downstream systems and be repeated across vast numbers of products. One well-designed systems feature can help customers obtain more useful work from GPUs already deployed while informing how future CUDA capabilities and GPU architectures are designed.
We are looking for systems software engineers who want to work at this leverage point. You will design and ship production C/C++ features and optimizations in the CUDA driver and runtime, trace important workloads across application, operating-system, CPU, interconnect, and GPU boundaries, bring up new platforms, and turn evidence into future software and hardware direction. Your work will not end at a benchmark: it can make AI tools more responsive and efficient, help scientists reach answers sooner, and enable intelligent machines and interactive products to operate within demanding real-time constraints. Over time, you can grow from owning critical features and performance paths to setting subsystem direction and leading hardware/software co-design across generations-helping build the computing foundation for the next decade of AI and accelerated computing.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD.
You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until September 5, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.