Senior Software Engineer - AI Inference Performance

NVIDIA AI

Santa Clara (CA)

On-site

USD 180,000 - 300,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Generous Benefits Package

Job summary

Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA.

Qualifications

  • Requires 6+ years in full-stack AI inference performance.
  • Strong programming skills in Python, C++, or Rust and expert CUDA.
  • Degree in Computer Science or Computer Engineering required.
  • Deep knowledge of GPU architecture and profiling tools.

Responsibilities

  • Lead end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems.
  • Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Skills

Full-stack AI
Python
C++
Rust
CUDA
GPU Architecture
Profiling Tools
Nsight Systems
Nsight Compute
PyTorch Profiler
TensorRT-LLM
vLLM
Triton
CUTLASS
Distributed Systems
Quantization

Education

Bachelor's degree in Computer Science or Computer Engineering

Tools

Nsight Systems
Nsight Compute
PyTorch Profiler
TensorRT-LLM
vLLM
Triton
CUTLASS

Job description

Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Requirements: Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA. A degree in Computer Science or Computer Engineering is required along with deep knowledge of GPU architecture and profiling tools.

Key Skills: CUDA, Python, C++, Rust, LLM Inference, VLM Inference, GPU Architecture, Nsight Systems, Nsight Compute, PyTorch Profiler, TensorRT-LLM, vLLM, Triton, CUTLASS, Distributed Systems, Quantization

Benefits: Equity, Generous Benefits Package

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Engineer (CUDA/LLM/VLM)
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA AI • Washington

On-site
USD 140,000 - 230,000
Equity
Comprehensive benefits package
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits