Software Engineer, CUDA Deep Learning Systems

Jobtailor

California (MO)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a highly skilled CUDA-focused engineer to optimize deep learning workloads across single-node to cluster-scale systems.

You will design custom CUDA kernels, improve performance for transformer-based models, and collaborate with researchers and kernel experts to push the boundaries of AI acceleration.

Qualifications

  • BS/MS/PhD in CS/CE/EE or related field, or equivalent experience.
  • 2+ years of industry or academic experience after degree.
  • Proficiency in C++ and Python.
  • Strong fundamentals in deep learning, especially transformers.
  • Understanding of distributed computing and multi-node scaling.
  • Experience in systems programming and low-level performance optimization.
  • Familiarity with GPU DL accelerator architectures.
  • Hands-on CUDA programming, kernel optimization, and profiling.
  • Experience with generative AI models and LLMs.
  • Research background in ML systems or related fields.

Responsibilities

  • Explore, research, and prototype systems optimizations for deep learning models and CUDA.
  • Architect and optimize distributed computing systems from single-node to cluster-scale environments.
  • Design and optimize custom high-performance CUDA kernels for new neural architectures.
  • Analyze hardware-software interactions to resolve bottlenecks in training and inference.
  • Collaborate with researchers, architects, kernel authors, and CUDA experts to co-design systems.
  • Develop tools and runtimes to profile and accelerate new deep learning paradigms.
  • Write clean, maintainable code and transition prototypes into open-source releases or products.

Skills

CUDA Programming
C++
Python
Deep Learning Fundamentals
Distributed Computing
Computer Architecture
Kernel Optimization
Workload Profiling
Generative AI Models
Transformers

Education

BS/MS/PhD in CS/CE/EE or equivalent

Tools

PyTorch
JAX
TensorRT
NCCL
MPI
UCX
Triton
XLA
Torch.compile
SgLang

Job description

  • Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA.
  • Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments.
  • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
  • Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products.
Requirements
  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
  • 2+ years of relevant industry experience or equivalent academic experience after degree achievement.
  • Strong proficiency in C++ and Python programming.
  • Solid background in deep learning fundamentals, focused on transformers.
  • Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with GPU deep learning accelerator architectures.
  • Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
  • Experience profiling and optimizing generative AI models, including large language models.
  • Research background in machine learning systems or adjacent fields.
  • Experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models.
  • Track record of initiative and willingness to deep-dive on problems across the stack.
  • Preferred experience with PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron internals and execution graphs.
  • Preferred hands-on experience with NCCL, MPI, or UCX and distributed machine learning techniques.
  • Preferred knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, and INT8.
  • Preferred background in deep learning compilers and ML systems, including Triton, XLA, and torch.compile.
  • Preferred experience with highly parallel or reinforcement-learning-style simulation environments.
  • Preferred experience designing and implementing agentic AI systems for complex systems and infrastructure problems.
Core Competencies

Demonstrates expertise in CUDA programming, high-performance computing, and deep learning optimization, with a strong foundation in distributed systems and machine learning architectures. Capable of collaborating with cross-functional teams to design and implement innovative solutions for advanced AI models.

Highest-signal resume keywords
  • CUDA Programming
  • C++ and Python Proficiency
  • Deep Learning Optimization
  • Distributed Computing Systems
  • Performance Bottleneck Analysis
ATS Optimization Keywords
Hard Skills
  • CUDA
  • C++
  • Python
  • Deep Learning Fundamentals
  • Systems Programming
  • Computer Architecture
  • Kernel Optimization
  • Workload Profiling
  • Generative AI Models
  • Transformers
Soft Skills
  • Collaboration
  • Initiative
  • Problem-Solving
Industry Keywords
  • Machine Learning Systems
  • High-Performance Computing
  • AI Research
  • Neural Network Architectures
  • Cluster-Scale Performance
Tools & Technologies
  • PyTorch
  • JAX
  • TensorRT
  • NCCL
  • MPI
  • UCX
  • Triton
  • XLA
  • Torch.compile
  • SgLang
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Performance Engineer – DGX Cloud
Senior Performance Engineer – DGX Cloud

Jobtailor • California (MO)

On-site
USD 180,000 - 280,000
Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Senior AI Systems and Algorithms Engineer
Senior AI Systems and Algorithms Engineer

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • New York (NY)

On-site
USD 180,000 - 260,000
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
System Software Engineer – AI
System Software Engineer – AI

Jobtailor • California (MO)

On-site
USD 120,000 - 170,000
Software Engineer, Neural Graphics Developer Tools
Software Engineer, Neural Graphics Developer Tools

Jobtailor • California (MO)

On-site
USD 140,000 - 230,000