Senior ML Inference Engineer – TensorRT & LLMs

NVIDIA

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking talented engineers for the TensorRT team to advance the industry-leading deep learning inference software. You will design, develop and optimize TensorRT/TensorRT-LLM for diverse inference workloads across data centers, workstations and PCs.

The role requires strong C++, Python, and CUDA skills, plus experience with DL frameworks and compilers. You will collaborate with DL experts and GPU architects to influence HW/SW design in a fast-paced environment.

Qualifications

  • BS/MS/PhD or equivalent experience in Computer Science or related field.
  • 4+ years of software development on a large codebase or project.
  • Strong proficiency in C++ (required), Rust or Python.
  • Experience in developing Deep Learning frameworks, compilers, or system software.
  • Excellent problem-solving and collaborative skills.

Responsibilities

  • Design, develop and optimize TensorRT and TensorRT-LLM for inference applications.
  • Develop software in C++, Python, and CUDA for deploying state-of-the-art LLMs and Generative AI models.
  • Collaborate with DL experts and GPU architects to influence hardware and software design for inference.

Skills

C++
Rust
Python
CUDA

Education

BS in CS/CE
MS in CS/CE
PhD in CS/CE

Tools

TensorRT
PyTorch
JAX

Job description

NVIDIA is seeking talented engineers for the TensorRT team to advance the industry-leading deep learning inference software. You will design, develop and optimize TensorRT/TensorRT-LLM for diverse inference workloads across data centers, workstations and PCs.

The role requires strong C++, Python, and CUDA skills, plus experience with DL frameworks and compilers. You will collaborate with DL experts and GPU architects to influence HW/SW design in a fast-paced environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer - TensorRT/CUDA
Senior ML Inference Engineer - TensorRT/CUDA

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior ML Inference Engineer - GPUs & TensorRT
Senior ML Inference Engineer - GPUs & TensorRT

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, DL Inference (TensorRT)
Senior Software Engineer, DL Inference (TensorRT)

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Edge AI Engineer: LLM Inference (TensorRT)
Senior Edge AI Engineer: LLM Inference (TensorRT)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Hybrid work model
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Software Engineer, Deep Learning Inference - TensorRT
Senior Software Engineer, Deep Learning Inference - TensorRT

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits