Senior CUDA DL Systems Engineer - Performance & Kernels Equity

NVIDIA

California (MO)

On-site

USD 184,000 - 357,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits
Career growth

Job summary

NVIDIA is seeking an experienced software professional to work at the intersection of CUDA and Deep Learning Systems, exploring next-generation ideas that bridge deep learning and accelerator architectures.

Join a research-oriented team to optimize hardware performance for AI workloads, designing custom CUDA kernels, distributed systems, and profiling tools. This role emphasizes collaboration with AI researchers and HW/SW architects to enhance compute utilization across nodes and clusters.

Qualifications

  • 8+ years of relevant industry experience or equivalent academic experience.
  • Strong proficiency in C++ and Python programming.
  • Solid background in the fundamentals of Deep Learning with a focus on transformers.
  • Strong understanding of distributed computing principles, multi-node scaling.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with deep learning accelerator architectures such as the GPU and hands-on experience with CUDA programming, kernel optimization, and workload profiling.

Responsibilities

  • Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level DL frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
  • Architect and optimize distributed computing systems that scale from a single node to cluster-scale environments.
  • Design, implement, and optimize custom high-performance CUDA kernels tailored to emerging neural network architectures and workloads.
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • Collaborate with AI researchers, HW/SW architects, kernel and compiler authors, and CUDA driver experts to co-design compute-efficient systems.
  • Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning.
  • Write clean, maintainable code, ensuring prototypes can translate into open-source releases or commercial products.

Skills

C++
Python
Deep Learning
Distributed computing
CUDA

Education

BS/MS/PhD in CS/CE/EE or related field

Tools

NCCL
MPI
UCX
Triton
XLA

Job description

NVIDIA is seeking an experienced software professional to work at the intersection of CUDA and Deep Learning Systems, exploring next-generation ideas that bridge deep learning and accelerator architectures.

Join a research-oriented team to optimize hardware performance for AI workloads, designing custom CUDA kernels, distributed systems, and profiling tools. This role emphasizes collaboration with AI researchers and HW/SW architects to enhance compute utilization across nodes and clusters.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior CUDA DL Systems Engineer for Performance & Scale
Senior CUDA DL Systems Engineer for Performance & Scale

NVIDIA • Town of Texas (WI)

On-site
USD 224,000 - 357,000
Equity and benefits
Senior CUDA & Deep Learning Systems Engineer
Senior CUDA & Deep Learning Systems Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
CUDA Deep Learning Systems Architect
CUDA Deep Learning Systems Architect

NVIDIA • Austin (TX)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior CUDA Driver Engineer — GPU Software & Performance
Senior CUDA Driver Engineer — GPU Software & Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Equity
Comprehensive benefits
Senior DL Performance Architect for Next-Gen AI Hardware
Senior DL Performance Architect for Next-Gen AI Hardware

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
CUDA System Software Engineer - Performance & GPU Kernel
CUDA System Software Engineer - Performance & GPU Kernel

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
Senior Software Engineer, CUDA Deep Learning Systems
Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Career growth
Senior Software Engineer, CUDA Deep Learning Systems
Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA • Town of Texas (WI)

On-site
USD 224,000 - 357,000
Equity and benefits
Senior Software Engineer, CUDA Deep Learning Systems
Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA • Austin (TX)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior AI/HPC GPU Engineer - Performance & Systems Expert
Senior AI/HPC GPU Engineer - Performance & Systems Expert

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits