Senior CUDA DL Systems Engineer for Performance & Scale

NVIDIA

Town of Texas (WI)

On-site

USD 224,000 - 357,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA is seeking an experienced software professional to push the boundaries of CUDA and Deep Learning Systems. You will work on architectures bridging DL models and CUDA, from kernel optimization to cluster-scale AI workloads.

Join a research-oriented team focused on profiling, optimizing, and accelerating AI workloads across single GPUs to supercomputer clusters. Strong C++, Python, and deep learning fundamentals are essential for success.

Qualifications

  • BS/MS/PhD in CS/CE/EE or related field (or equivalent experience).
  • 8+ years of relevant industry or academic experience after degree.
  • Strong proficiency in C++ and Python programming.
  • Solid background in Deep Learning fundamentals with focus on transformers.
  • Strong understanding of distributed computing and multi-node scaling.
  • Experience in CUDA programming, kernel optimization, and workload profiling.
  • Research background in ML systems or related fields with vision/diffusion/AI models.

Responsibilities

  • Explore novel system optimizations for deep learning models at the intersection of DL and CUDA.
  • Architect and optimize distributed systems from single node to cluster-scale environments.
  • Design and optimize custom high-performance CUDA kernels for neural network architectures.
  • Analyze hardware-software interactions to troubleshoot performance bottlenecks in training and inference.
  • Collaborate with AI researchers, HW/SW architects, kernel/driver experts to improve utilization and memory bandwidth.
  • Develop tools and runtimes to profile and accelerate new DL paradigms.
  • Write clean, maintainable code and prepare prototypes for open-source or commercial integration.

Skills

C++
Python
Deep Learning
Transformers
Distributed Computing
CUDA
Kernel Optimization
Workload Profiling
Research Experience

Education

BS/MS/PhD in CS/CE/EE or related

Tools

NCCL
MPI
UCX
TensorRT
Triton
Torch Compile

Job description

NVIDIA is seeking an experienced software professional to push the boundaries of CUDA and Deep Learning Systems. You will work on architectures bridging DL models and CUDA, from kernel optimization to cluster-scale AI workloads.

Join a research-oriented team focused on profiling, optimizing, and accelerating AI workloads across single GPUs to supercomputer clusters. Strong C++, Python, and deep learning fundamentals are essential for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Deep Learning Systems Engineer — Optimize AI at Scale
CUDA Deep Learning Systems Engineer — Optimize AI at Scale

NVIDIA • Town of Texas (WI)

On-site
USD 124,000 - 196,000
Equity
Benefits
CUDA Deep Learning Systems Architect
CUDA Deep Learning Systems Architect

NVIDIA • Austin (TX)

On-site
USD 224,000 - 357,000
Equity
Benefits
AI & GPU Systems Engineer — Deep Learning Performance
AI & GPU Systems Engineer — Deep Learning Performance

NVIDIA • Santa Clara (CA)

On-site
USD 108,000 - 196,000
Equity
Benefits
Inclusive work environment
Senior Software Engineer, CUDA Deep Learning Systems
Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA • Town of Texas (WI)

On-site
USD 224,000 - 357,000
Equity and benefits
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA • Town of Texas (WI)

On-site
USD 124,000 - 196,000
Equity
Benefits
Senior Software Engineer, CUDA Deep Learning Systems
Senior Software Engineer, CUDA Deep Learning Systems

NVIDIA • Austin (TX)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior CUDA Driver Engineer — GPU Software & Performance
Senior CUDA Driver Engineer — GPU Software & Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Comprehensive benefits
CUDA Systems Engineer - Performance Driver & Runtime
CUDA Systems Engineer - Performance Driver & Runtime

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity
Benefits
Senior DL Inference Engineer — GPU Performance OpenSource
Senior DL Inference Engineer — GPU Performance OpenSource

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits package
Competitive salary
Senior AI Networking and Performance Engineer for LLMs
Senior AI Networking and Performance Engineer for LLMs

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits