Senior GPU ML Inference Engineer — Edge AI Platforms

NVIDIA

Westford (MA)

On-site

USD 224,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is building the software stack for fast LLM inference on edge AI hardware in Westford, MA. This role focuses on evaluating open-source inference frameworks and mapping architectures to NVIDIA GPUs to maximize throughput and minimize latency.

You will own validation workflows, develop recipes, and collaborate with the community and partners to solve hardware-specific inference challenges, with equity and benefits on offer.

Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference
  • Strong Python or C++ programming, software design, and software engineering skills.
  • Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) — you understand how thread blocks, memory hierarchy, and warp execution affect real-world performance
  • Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism
  • Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit
  • Strong analytical skills: ability to form a performance hypothesis, design an experiment, interpret results, and communicate findings clearly

Responsibilities

  • Track and evaluate innovations in open-source LLM inference frameworks for NVIDIA edge AI hardware.
  • Analyze model architectures and inference algorithms on NVIDIA GPUs to identify mismatches and optimization opportunities.
  • Characterize multi-node inference: NCCL/RCCL, topology-aware all-reduce, and parallelism on edge clusters.
  • Produce performance reports mapping hardware limits to observed throughput, latency, and utilization.
  • Own model validation for new releases: architecture compatibility, recipe development, performance characterization, and publication to recipe sites.
  • Develop and maintain developer-facing inference recipes; automate staleness detection and CI feedback loops.
  • Engage with community and partners on model bring-up; be the technical contact for hardware-specific inference issues.

Skills

Python
C++
GPU computing
ML systems
Performance optimization
Triton
CUDA
KV-cache

Education

BS/MS/PhD in CS/CE/EE

Tools

NVIDIA Container Toolkit
Docker
OCI builds

Job description

NVIDIA is building the software stack for fast LLM inference on edge AI hardware in Westford, MA. This role focuses on evaluating open-source inference frameworks and mapping architectures to NVIDIA GPUs to maximize throughput and minimize latency.

You will own validation workflows, develop recipes, and collaborate with the community and partners to solve hardware-specific inference challenges, with equity and benefits on offer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU AI Platform Engineer — Edge Inference (Equity)
Senior GPU AI Platform Engineer — Edge Inference (Equity)

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior GPU AI Platforms Engineer - Edge LLM Inference
Senior GPU AI Platforms Engineer - Edge LLM Inference

NVIDIA • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
GPU Inference Engineer — Deep Learning
GPU Inference Engineer — Deep Learning

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Compiler Engineer - MLIR, GPU Inference, Equity
Senior AI Compiler Engineer - MLIR, GPU Inference, Equity

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Senior AI Systems Engineer – Edge GPU & Local Inference
Senior AI Systems Engineer – Edge GPU & Local Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Compiler Engineer — MLIR for GPU Inference
Senior AI Compiler Engineer — MLIR for GPU Inference

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Edge AI Inference Engineer
Senior Edge AI Inference Engineer

Intel • Santa Clara (CA)

Hybrid
USD 195,000 - 362,000