Senior Deep Learning Architect, LLM Inference

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

10 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Santa Clara, CA, seeks a Senior Deep Learning Architect specializing in LLM inference. You will characterize workloads on the latest LLMs and inference servers, comparing with vLLM, SGLang, and TRT-LLM to maintain NVIDIA's leadership in GPU-accelerated AI.

You will collaborate with engineers from AI startups, publish benchmarks, and develop profiling tools. A Level 4/5 base salary is offered, with equity and benefits; applications close Aug 15, 2026.

Qualifications

  • Master's or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 6+ years of relevant software development experience.
  • Deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client-server LLM applications with OpenAI API or MCP and identifying bottlenecks.
  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with large software projects like frameworks, compilers, or OS.
  • Proficient with AI coding agents like Claude Code, Codex, and Cursor.
  • Excellent written and verbal communication; able to work independently and collaboratively.

Responsibilities

  • Characterize workloads of the latest LLMs and inference servers to keep NVIDIA ahead.
  • Collaborate with performance marketing to publish blogs and updates to InferenceX.
  • Work with AI startup engineers to establish standard benchmarking methods.
  • Develop an evolving inference performance data website.
  • Create E2E profiling and analysis tools for rapid AI advances.
  • Contribute to PyTorch, TRT-LLM, vLLM, and SGLang projects to advance the field.
  • Verify that new GPU launches deliver industry-leading performance.
  • Guide inference serving direction across software, research, and product teams.

Skills

Deep Learning
PyTorch
Profiling
Performance optimization
OpenAI API experience
C++/Python

Education

Master's or PhD in CS/CE

Tools

vLLM
TRT-LLM
SGLang

Job description

We are now looking for a Senior Deep Learning Architect, LLM Inference!

NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS, then this is a great role for you!

What You’ll Be Doing
  • You will do workload characterization of the latest LLMs and inference servers like vLLM, SGLang and TRT-LLM to ensure NVIDIA maintains its leadership position.
  • Join forces with the performance marketing team to build engaging content, including blog posts and updates to InferenceX to highlight NVIDIA's outstanding inference achievements.
  • Collaborate with engineers from AI startup companies to establish standard benchmarking methodologies.
  • Develop a constantly evolving inference performance data results website.
  • Invent E2E profiling and analysis tools that you will use to keep up with the rapid pace of Generative AI.
  • Contribute to deep learning software projects, such as PyTorch, TRT-LLM, vLLM, and SGLang to drive advancements in the field.
  • Verify that new GPU product launches produce industry leading performance.
  • Collaborate across the company to guide the direction of inference serving, working with software, research, and product teams to ensure best-in-class performance.
  • Use the latest coding agents and inference technology to improve team efficiency.
What We Need To See
  • Master's or PhD degree in Computer Science, Computer Engineering, related fields, or equivalent experience.
  • 6+ years of relevant software development experience.
  • Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.
  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with complex software projects like frameworks, compilers, or operating systems.
  • Demonstrated proficiency with the latest AI coding agents like Claude Code, Codex, and Cursor
  • Excellent written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.
Ways To Stand Out From The Crowd
  • Demonstrate a drive to continuously improve software and hardware performance.
  • Showcase examples of novel use cases for agentic AI tools in the workplace.
  • Experience with databases and visualization tools will set you apart.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have a team of highly skilled and motivated individuals who excel in their work. If you have a proactive and independent approach, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 15, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits package
Competitive salary
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity opportunities
Comprehensive benefits package
Senior Research Engineer - Enterprise Products
Senior Research Engineer - Enterprise Products

NVIDIA Gruppe • Washington

On-site
USD 192,000 - 356,500
Equity
Benefits