Senior DL Engineer — Inference & Model Optimization

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package
Competitive salary

Job summary

NVIDIA is seeking a Senior Deep Learning Software Engineer to scale automated inference and deployment for generative AI models in Santa Clara, CA. You will train and deploy state-of-the-art LLMs and diffusion models, using the Torch 2.0 stack to extract graph representations and optimize inference with advanced parallelism and kernels.

Work across teams to push performance, profiling GPU kernels and refining deployment strategies for TRT and related tools.

Qualifications

  • Masters, PhD, or equivalent in CS/AI or related field.
  • 8+ years of deep learning experience.
  • Strong software design, debugging, and performance analysis.
  • Proficiency in Python, PyTorch, and ML tools (e.g. HuggingFace).
  • Strong algorithms and programming fundamentals.
  • Excellent written and verbal communication; able to work independently and with teams.

Responsibilities

  • Train, develop, and deploy state-of-the-art generative AI models like LLMs and diffusion models.
  • Leverage Torch 2.0 ecosystem to analyze models and extract graph representations.
  • Develop high-performance inference optimization techniques and tensor parallelism.
  • Collaborate across teams to implement efficient kernel implementations.
  • Profile GPU kernel performance and identify optimization opportunities.
  • Innovate on inference performance to maintain leadership in NVIDIA's software stack.
  • Architect modular, scalable software platforms with broad model support.

Skills

Deep Learning
Python
PyTorch
Algorithms
Software Design
Communication
HuggingFace

Education

Masters or PhD in CS/AI or related field

Tools

CUDA
Triton
CUTLASS

Job description

NVIDIA is seeking a Senior Deep Learning Software Engineer to scale automated inference and deployment for generative AI models in Santa Clara, CA. You will train and deploy state-of-the-art LLMs and diffusion models, using the Torch 2.0 stack to extract graph representations and optimize inference with advanced parallelism and kernels.

Work across teams to push performance, profiling GPU kernels and refining deployment strategies for TRT and related tools.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DL Engineer – Inference & Model Optimization (Equity)
Senior DL Engineer – Inference & Model Optimization (Equity)

NVIDIA • United States

Remote
USD 184,000 - 356,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior DL Inference Engineer - GPU & LLM Performance
Senior DL Inference Engineer - GPU & LLM Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer — GPU Performance OpenSource
Senior DL Inference Engineer — GPU Performance OpenSource

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits package
Competitive salary
Engineering Manager, AI Inference & GPU Scaling
Engineering Manager, AI Inference & GPU Scaling

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model
Senior Inference Engineer: AI-Driven GPU Kernel Optimization
Senior Inference Engineer: AI-Driven GPU Kernel Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior Deep Learning Software Engineer, Inference and Model Optimization
Senior Deep Learning Software Engineer, Inference and Model Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Competitive salary
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA AI • Washington

On-site
USD 140,000 - 230,000
Equity
Comprehensive benefits package