Senior DL Software Engineer - Inference & Model Optimization (Equity)

2100 NVIDIA USA

Santa Clara (CA)

On-site

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

NVIDIA in Santa Clara, CA, is seeking a Senior Deep Learning Software Engineer to scale our automated inference and deployment solution within the Algorithmic Model Optimization Team. You will work across the ML stack, from PyTorch and HuggingFace to high‑performance CUDA kernels and the TRT Model Optimizer, building a modular, scalable platform.

Applicants should have a Masters or PhD (or equivalent), 5+ years in deep learning, excellent Python and PyTorch skills, and strong software design and

Qualifications

  • Masters, PhD, or equivalent experience in Computer Science, AI, Applied Math, or related field.
  • 5+ years of relevant work or research experience in Deep Learning.
  • Excellent software design skills, including debugging, performance analysis, and test design.
  • Strong proficiency in Python, PyTorch, and related ML tools (e.g. HuggingFace).
  • Strong algorithms and programming fundamentals.
  • Good written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.

Responsibilities

  • Train, develop, and deploy state‑of‑the generative AI models like LLMs and diffusion models using NVIDIA's AI software stack.
  • Leverage Torch 2.0 ecosystem to analyze and extract standardized model graphs for automated deployment.
  • Develop high-performance optimization techniques for inference (e.g., tensor parallelism, kv-caching).
  • Collaborate with teams across NVIDIA to implement performant kernel code in our deployment solution.
  • Analyze and profile GPU kernel-level performance to identify optimization opportunities.
  • Innovate on inference performance to maintain leadership with TRT, TRT‑LLM, and TRT Model Optimizer.
  • Architect and design a modular, scalable platform with broad model support and optimization techniques.

Skills

Python
PyTorch
HuggingFace
Algorithms
Software design

Education

Masters or PhD in CS/AI/Applied Math

Tools

CUDA
TensorRT
Triton
Cutlass

Job description

NVIDIA in Santa Clara, CA, is seeking a Senior Deep Learning Software Engineer to scale our automated inference and deployment solution within the Algorithmic Model Optimization Team. You will work across the ML stack, from PyTorch and HuggingFace to high‑performance CUDA kernels and the TRT Model Optimizer, building a modular, scalable platform.

Applicants should have a Masters or PhD (or equivalent), 5+ years in deep learning, excellent Python and PyTorch skills, and strong software design and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DL Inference Engineer – GPU Optimization Equity
Senior DL Inference Engineer – GPU Optimization Equity

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity options
Comprehensive benefits
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference and Model Optimization
Senior Deep Learning Software Engineer, Inference and Model Optimization

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Software Engineer, DL Inference (TensorRT)
Senior Software Engineer, DL Inference (TensorRT)

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Hardware Modeling Architect — Equity Options
Senior DL Hardware Modeling Architect — Equity Options

NVIDIA • Massachusetts

On-site
USD 152,000 - 288,000
Senior LLM Inference Architect — Equity Eligible, Remote
Senior LLM Inference Architect — Equity Eligible, Remote

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Senior ML Inference Engineer - GPUs & TensorRT
Senior ML Inference Engineer - GPUs & TensorRT

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior CUDA Deep Learning Systems Engineer – Equity
Senior CUDA Deep Learning Systems Engineer – Equity

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package