System Software Engineer, Performance - CUDA Driver

NVIDIA

United States

On-site

USD 130,000 - 170,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking systems software engineers to design and ship production C/C++ features in the CUDA driver and runtime, tracing workloads from application to GPU. You will optimize critical paths, bring up new platforms, and translate evidence into future software directions, collaborating across teams to improve performance and scalability.

Ideal candidates have strong OS/concurrency foundations, solid computer-architecture knowledge, and a track record of delivering production performance

Qualifications

  • BS/MS/PhD or equivalent practical experience in CS/CE/EE with 2+ years in systems software.
  • Strong production C/C++ systems programming and production features.
  • Solid OS and concurrency fundamentals; hardware and memory systems awareness.

Responsibilities

  • Design, implement, validate, and ship performance-centric CUDA driver/runtime features in production C/C++.
  • Trace workloads across application, OS, CPU, interconnect, and GPU boundaries.
  • Bring up new platforms and turn evidence into software and hardware direction.
  • Own complex performance problems end-to-end across software/hardware boundaries.
  • Set performance expectations and drive readiness for future silicon and platforms.
  • Translate workload evidence into CUDA API and programming-model improvements.
  • Collaborate with cross-functional teams; communicate findings and improve engineering quality.

Skills

Production C/C++ systems-programming
Operating systems & concurrency
Computer architecture foundations
Performance optimization
Communication & collaboration
CUDA/GPU experience (beneficial)
Ownership of complex problems

Education

BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering or equivalent

Tools

CUDA

Job description

The AI revolution is not powered by models alone, rather it advances when enormous amounts of computation become fast, efficient, and economical enough to turn new ideas into products people can use on a global scale. Faster training lets research and product teams test the next idea sooner. Lower-latency, higher-throughput inference makes AI assistants and agents more responsive and practical for more people. Shorter time to solution lets scientists and engineers explore more possibilities within the same time and energy budget.

At NVIDIA, performance is not a supporting metric - it is how architectural invention becomes useful computing. CUDA is a critical layer where that transformation happens, sitting beneath the frameworks, libraries, and applications used across AI, deep learning, and HPC, as well as graphics, automotive, robotics, and other CUDA-powered products. That gives this team unusual leverage: reduce overhead in a fundamental launch, synchronization, memory, or data-movement path-or create a new driver or runtime capability-and the improvement can flow through many downstream systems and be repeated across vast numbers of products. One well-designed systems feature can help customers obtain more useful work from GPUs already deployed while informing how future CUDA capabilities and GPU architectures are designed.

We are looking for systems software engineers who want to work at this leverage point. You will design and ship production C/C++ features and optimizations in the CUDA driver and runtime, trace important workloads across application, operating-system, CPU, interconnect, and GPU boundaries, bring up new platforms, and turn evidence into future software and hardware direction. Your work will not end at a benchmark: it can make AI tools more responsive and efficient, help scientists reach answers sooner, and enable intelligent machines and interactive products to operate within demanding real-time constraints. Over time, you can grow from owning critical features and performance paths to setting subsystem direction and leading hardware/software co-design across generations-helping build the computing foundation for the next decade of AI and accelerated computing.

What you’ll be doing:
  • Design, implement, validate, and ship performance-centric features and programming-model capabilities in the CUDA driver and runtime, writing maintainable, well-tested production C/C++.
  • Optimize critical execution paths-including kernel launch, synchronization, memory management & movement, CPU-GPU coordination, and system interconnect use-for latency, throughput, bandwidth, efficiency, and scalability.
  • Own complex performance problems end-to-end - understand important workloads, form hypotheses, create focused measurements and models, isolate root causes across software and hardware boundaries, implement production solutions, and validate application-level impact.
  • Establish performance expectations for current and future platforms, characterize new silicon, close software and hardware gaps, and drive performance readiness through product release.
  • Translate workload and platform evidence into CUDA API and programming-model improvements, systems-software direction, and measurement-backed recommendations for future hardware architecture and implementation.
  • Partner with application, library, framework, operating-system, driver, runtime, firmware, GPU architecture, silicon, product, and customer-facing teams; communicate findings clear and raise engineering quality through design and code reviews.
  • Lead complex feature development and cross-layer investigations across teams, define performance requirements and technical direction for major subsystems, mentor engineers, and shape hardware/software decisions for future product generations.
What we need to see:
  • A BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field-or equivalent practical experience - with at least 2 years of relevant systems-software development experience.
  • Strong production C/C++ systems-programming experience, including delivery of substantial features, optimizations, or production fixes in a complex codebase.
  • Strong operating systems and concurrency foundations, including threads, synchronization, processes, virtual memory, and user/kernel interactions.
  • Strong computer-architecture foundations, including processors, memory hierarchy, caching and coherence, data movement, and system interconnects.
  • Demonstrated success improving real software performance: measuring behavior, identifying the limiting mechanism, implementing an effective solution, and validating the result quantitatively.
  • Sound technical judgment, ownership of ambiguous problems, and clear communication across organizational and disciplinary boundaries.
  • Direct CUDA or GPU experience is valuable but is not required when accompanied by deep systems-software, operating-systems, computer-architecture,
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Software Engineer, Performance - CUDA Driver
System Software Engineer, Performance - CUDA Driver

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits
System Software Engineer, Performance - CUDA Driver
System Software Engineer, Performance - CUDA Driver

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 124,000 - 196,000
System Software Engineer, Performance - CUDA Driver
System Software Engineer, Performance - CUDA Driver

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 124,000 - 196,000
System Software Engineer, Performance - CUDA Driver
System Software Engineer, Performance - CUDA Driver

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity
Benefits
System Software Engineer, Performance - CUDA Driver
System Software Engineer, Performance - CUDA Driver

Socket.dev • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits
System Software Engineer, Performance - CUDA Driver
System Software Engineer, Performance - CUDA Driver

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
CUDA Systems Engineer - Performance Driver & Runtime
CUDA Systems Engineer - Performance Driver & Runtime

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity
Benefits
CUDA Systems Engineer - Performance & Runtime
CUDA Systems Engineer - Performance & Runtime

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 124,000 - 196,000
CUDA System Software Engineer - Performance & GPU Kernel
CUDA System Software Engineer - Performance & GPU Kernel

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
Performance Systems Engineer, CUDA Driver & Runtime
Performance Systems Engineer, CUDA Driver & Runtime

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 124,000 - 196,000