Senior Deep Learning Architect, LLM Inference

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA in Santa Clara is seeking a seasoned software professional to characterize workloads for the latest LLMs and inference servers. You will join forces with the performance marketing team to craft engaging content and showcase NVIDIA's inference achievements, while collaborating with AI startup engineers to set benchmarking standards.

You will help build a dynamic inference data results site and develop end-to-end profiling tools to keep pace with Generative AI advancements, contributing to

Qualifications

  • Master's or PhD in Computer Science, Computer Engineering, or related fields, or equivalent experience.
  • 6+ years of software development experience.
  • Deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client-server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.
  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with complex software projects like frameworks, compilers, or operating systems.
  • Proficiency with AI coding agents like Claude Code, Codex, Cursor.
  • Excellent written and verbal communication skills and ability to work independently and collaboratively in a fast-paced environment.

Responsibilities

  • Characterize workload of LLMs and inference servers to maintain leadership.
  • Collaborate with performance marketing to create content and updates.
  • Work with engineers from AI startups to establish benchmarking methodologies.
  • Develop an evolving inference performance data results website.
  • Create E2E profiling and analysis tools to keep pace with Generative AI.
  • Contribute to PyTorch, TRT-LLM, vLLM, and SGLang to drive AI advancements.
  • Verify that new GPU launches deliver industry-leading performance.
  • Guide inference serving direction across software, research, and product teams.
  • Utilize the latest coding agents and inference tech to improve team efficiency.

Skills

DL inference serving
PyTorch
Profiling
Compiler optimizations
OpenAI API
MCP
CPU/GPU microarchitecture
LLM apps
AI agents
Communication

Education

Master's or PhD in CS/CE

Tools

Claude Code
Codex
Cursor

Job description

What you\'ll be doing:


  • You will do workload characterization of the latest LLMs and inference servers like vLLM, SGLang and TRT-LLM to ensure NVIDIA maintains its leadership position.

  • Join forces with the performance marketing team to build engaging content, including blog posts and updates to InferenceX to highlight NVIDIA\'s outstanding inference achievements.

  • Collaborate with engineers from AI startup companies to establish standard benchmarking methodologies.

  • Develop a constantly evolving inference performance data results website.

  • Invent E2E profiling and analysis tools that you will use to keep up with the rapid pace of Generative AI.

  • Contribute to deep learning software projects, such as PyTorch, TRT-LLM, vLLM, and SGLang to drive advancements in the field.

  • Verify that new GPU product launches produce industry leading performance.

  • Collaborate across the company to guide the direction of inference serving, working with software, research, and product teams to ensure best-in-class performance.

  • Use the latest coding agents and inference technology to improve team efficiency.



What we need to see:



  • Master\'s or PhD degree in Computer Science, Computer Engineering, related fields, or equivalent experience.

  • 6+ years of relevant software development experience.

  • Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.

  • Experience developing client server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.

  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.

  • Experience with complex software projects like frameworks, compilers, or operating systems.

  • Demonstrated proficiency with the latest AI coding agents like Claude Code, Codex, and Cursor

  • Excellent written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.



Ways to stand out from the crowd:



  • Demonstrate a drive to continuously improve software and hardware performance.

  • Showcase examples of novel use cases for agentic AI tools in the workplace.

  • Experience with databases and visualization tools will set you apart.


NVIDIA is widely considered to be one of the technology world\'s most desirable employers. We have a team of highly skilled and motivated individuals who excel in their work. If you have a proactive and independent approach, we want to hear from you!


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.


You will also be eligible for equity and benefits.


Applications for this job will be accepted at least until August 15, 2026.


This posting is for an existing vacancy.


NVIDIA pledges to foster an inclusive work environment and proudly
to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Principal Deep Learning Algorithm Engineer
Principal Deep Learning Algorithm Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Socket.dev • Santa Clara (UT)

On-site
USD 224,000 - 357,000
Equity
Benefits package
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Deep Learning Software Engineer, Inference - New College Grad 2026
Deep Learning Software Engineer, Inference - New College Grad 2026

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior System Software Engineer, Agentic Inference - Dynamo
Senior System Software Engineer, Agentic Inference - Dynamo

NVIDIA • Santa Clara (CA)

Hybrid
USD 272,000 - 431,250
Equity
Benefits
Hybrid work
Principal Deep Learning Algorithm Engineer
Principal Deep Learning Algorithm Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits