Senior LLM Inference Architect — Performance & Benchmarking

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA in Santa Clara is seeking a seasoned software professional to characterize workloads for the latest LLMs and inference servers. You will join forces with the performance marketing team to craft engaging content and showcase NVIDIA's inference achievements, while collaborating with AI startup engineers to set benchmarking standards.

You will help build a dynamic inference data results site and develop end-to-end profiling tools to keep pace with Generative AI advancements, contributing to

Qualifications

  • Master's or PhD in Computer Science, Computer Engineering, or related fields, or equivalent experience.
  • 6+ years of software development experience.
  • Deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client-server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.
  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with complex software projects like frameworks, compilers, or operating systems.
  • Proficiency with AI coding agents like Claude Code, Codex, Cursor.
  • Excellent written and verbal communication skills and ability to work independently and collaboratively in a fast-paced environment.

Responsibilities

  • Characterize workload of LLMs and inference servers to maintain leadership.
  • Collaborate with performance marketing to create content and updates.
  • Work with engineers from AI startups to establish benchmarking methodologies.
  • Develop an evolving inference performance data results website.
  • Create E2E profiling and analysis tools to keep pace with Generative AI.
  • Contribute to PyTorch, TRT-LLM, vLLM, and SGLang to drive AI advancements.
  • Verify that new GPU launches deliver industry-leading performance.
  • Guide inference serving direction across software, research, and product teams.
  • Utilize the latest coding agents and inference tech to improve team efficiency.

Skills

DL inference serving
PyTorch
Profiling
Compiler optimizations
OpenAI API
MCP
CPU/GPU microarchitecture
LLM apps
AI agents
Communication

Education

Master's or PhD in CS/CE

Tools

Claude Code
Codex
Cursor

Job description

NVIDIA in Santa Clara is seeking a seasoned software professional to characterize workloads for the latest LLMs and inference servers. You will join forces with the performance marketing team to craft engaging content and showcase NVIDIA's inference achievements, while collaborating with AI startup engineers to set benchmarking standards.

You will help build a dynamic inference data results site and develop end-to-end profiling tools to keep pace with Generative AI advancements, contributing to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior LLM Inference Architect
Senior LLM Inference Architect

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Platform Product Lead
Senior AI Inference Platform Product Lead

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Comprehensive benefits
Inclusive culture
Senior LLM Inference Architect — Equity Eligible, Remote
Senior LLM Inference Architect — Equity Eligible, Remote

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Distributed AI Inference Performance Engineer
Distributed AI Inference Performance Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Technical Lead: Inference Benchmarking & ML Infra
Technical Lead: Inference Benchmarking & ML Infra

NVIDIA • United States

On-site
USD 224,000 - 357,000
Equity and benefits
Comprehensive benefits package
Competitive salaries