Senior ML Inference Engineer - GPUs & TensorRT

NVIDIA AI

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in California is seeking an experienced Senior Software Engineer to join the TensorRT team, shaping the next generation of inference software for AI accelerators. You will design and optimize TensorRT and TensorRT-LLM, write in C++, Python and CUDA, and work with GPU architects to push performance on modern GPUs.

This role requires advanced degrees and 4+ years of experience, offers equity and benefits, and a clear path to impactful work on real-time AI inference.

Qualifications

  • BS, MS, PhD or equivalent in Computer Science, Computer Engineering or related field.
  • 4+ years of software development experience on a large codebase or project.
  • Strong proficiency in C++ (required), Rust or Python programming languages.
  • Experience in developing Deep Learning Frameworks, Compilers, or System Software.
  • Excellent problem-solving skills and passion to learn and work effectively in a fast-paced, collaborative environment.
  • Strong communication skills and the ability to articulate complex technical concepts.

Responsibilities

  • Design, develop and optimize NVIDIA TensorRT and TensorRT-LLM to supercharge inference applications for datacenter, workstations, and PCs.
  • Develop software in C++, Python, and CUDA for seamless and efficient deployment of state-of-the-art LLMs and Generative AI models.
  • Collaborate with deep learning experts and GPU architects throughout the company to influence Hardware and Software design for inference.

Skills

C++
Python
Rust
CUDA
Deep Learning Frameworks

Education

BS/MS/PhD in Computer Science or Computer Engineering or related field

Tools

TensorRT

Job description

NVIDIA in California is seeking an experienced Senior Software Engineer to join the TensorRT team, shaping the next generation of inference software for AI accelerators. You will design and optimize TensorRT and TensorRT-LLM, write in C++, Python and CUDA, and work with GPU architects to push performance on modern GPUs.

This role requires advanced degrees and 4+ years of experience, offers equity and benefits, and a clear path to impactful work on real-time AI inference.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer - TensorRT/CUDA
Senior ML Inference Engineer - TensorRT/CUDA

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior ML Inference Engineer – TensorRT & LLMs
Senior ML Inference Engineer – TensorRT & LLMs

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, DL Inference (TensorRT)
Senior Software Engineer, DL Inference (TensorRT)

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
High-Performance AI Inference Engineer (TensorRT)
High-Performance AI Inference Engineer (TensorRT)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Deep Learning Inference - TensorRT
Senior Software Engineer, Deep Learning Inference - TensorRT

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model