Senior LLM Inference Architect - Equity Eligible

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

18 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Santa Clara, CA, seeks a Senior Deep Learning Architect specializing in LLM inference. You will characterize workloads on the latest LLMs and inference servers, comparing with vLLM, SGLang, and TRT-LLM to maintain NVIDIA's leadership in GPU-accelerated AI.

You will collaborate with engineers from AI startups, publish benchmarks, and develop profiling tools. A Level 4/5 base salary is offered, with equity and benefits; applications close Aug 15, 2026.

Qualifications

  • Master's or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 6+ years of relevant software development experience.
  • Deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client-server LLM applications with OpenAI API or MCP and identifying bottlenecks.
  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with large software projects like frameworks, compilers, or OS.
  • Proficient with AI coding agents like Claude Code, Codex, and Cursor.
  • Excellent written and verbal communication; able to work independently and collaboratively.

Responsibilities

  • Characterize workloads of the latest LLMs and inference servers to keep NVIDIA ahead.
  • Collaborate with performance marketing to publish blogs and updates to InferenceX.
  • Work with AI startup engineers to establish standard benchmarking methods.
  • Develop an evolving inference performance data website.
  • Create E2E profiling and analysis tools for rapid AI advances.
  • Contribute to PyTorch, TRT-LLM, vLLM, and SGLang projects to advance the field.
  • Verify that new GPU launches deliver industry-leading performance.
  • Guide inference serving direction across software, research, and product teams.

Skills

Deep Learning
PyTorch
Profiling
Performance optimization
OpenAI API experience
C++/Python

Education

Master's or PhD in CS/CE

Tools

vLLM
TRT-LLM
SGLang

Job description

NVIDIA in Santa Clara, CA, seeks a Senior Deep Learning Architect specializing in LLM inference. You will characterize workloads on the latest LLMs and inference servers, comparing with vLLM, SGLang, and TRT-LLM to maintain NVIDIA's leadership in GPU-accelerated AI.

You will collaborate with engineers from AI startups, publish benchmarks, and develop profiling tools. A Level 4/5 base salary is offered, with equity and benefits; applications close Aug 15, 2026.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Deep Learning Architect, LLM Inference
Senior Deep Learning Architect, LLM Inference

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior LLM Infra Engineer — AI Model Serving
Senior LLM Infra Engineer — AI Model Serving

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits package
Competitive salary
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
LLM Efficiency Architect — Lead Cross‑Layer Optimization (Equity)
LLM Efficiency Architect — Lead Cross‑Layer Optimization (Equity)

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior DL Inference Engineer - GPU & LLM Performance
Senior DL Inference Engineer - GPU & LLM Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000