Senior AI Inference Architect — Disaggregated Serving

Nvidia Corporation in

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a Senior Solutions Architect for the Inference Service Providers team to redefine how AI inference scales with finite resources. You will lead high-stakes engagements with partners, design end-to-end reference architectures, and mentor teams across customers and internal groups.

You’ll master full-stack inference, optimize with tools like Dynamo, Triton, and TensorRT-LLM, and help customers deploy disaggregated inference at scale while evangelizing DevOps best practices for

Qualifications

  • 6+ years in solutions architecture driving customer engagements to deploy distributed systems, including 2+ years with AI workloads on Kubernetes.
  • Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization and serving.
  • Deep knowledge of modern inference best practices including disaggregated serving, multi-tier KV cache management, speculative decoding, quantization and custom inference kernels.
  • Hands-on full-stack agent design with modern sandboxing, memory and retrieval systems with evaluation, skill design and governance.
  • BS in CS/Engineering or equivalent experience.

Responsibilities

  • Work with inference partners to teach the value of our stack and bring back product insights.
  • Build and operate inference recipes with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency.
  • Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang to ensure seamless integration with disaggregated inference.
  • Evangelize DevOps best-practices for managing Kubernetes clusters, configuring compute fabrics, and observability.
  • Provide mentorship and technical leadership to customers and internal teams, guiding deployment of disaggregated inference systems and resolving complex issues.

Skills

Solutions Architecture
Kubernetes
TensorRT-LLM
NVIDIA Dynamo
vLLM
Inference pipelines

Education

BS in CS/Engineering

Tools

NVIDIA Dynamo
Triton Inference Server
TensorRT-LLM
vLLM

Job description

NVIDIA is seeking a Senior Solutions Architect for the Inference Service Providers team to redefine how AI inference scales with finite resources. You will lead high-stakes engagements with partners, design end-to-end reference architectures, and mentor teams across customers and internal groups.

You’ll master full-stack inference, optimize with tools like Dynamo, Triton, and TensorRT-LLM, and help customers deploy disaggregated inference at scale while evangelizing DevOps best practices for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Architect — Scale & Disaggregated Serving
Senior AI Inference Architect — Scale & Disaggregated Serving

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Lead Solutions Architect — AI Inference & Disaggregated Serving
Lead Solutions Architect — AI Inference & Disaggregated Serving

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Competitive salaries
Comprehensive benefits package
Equal opportunity employer
Lead Architect, AI Inference & Disaggregated Serving
Lead Architect, AI Inference & Disaggregated Serving

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Senior Architect, Scaled AI Inference & Systems — Equity
Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Senior Architect, Scaled AI Inference & Orchestration
Senior Architect, Scaled AI Inference & Orchestration

NVIDIA • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior Solutions Architect, Inference Service Providers
Senior Solutions Architect, Inference Service Providers

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Competitive salaries
Comprehensive benefits package
Equal opportunity employer
Senior Solutions Architect, AI Inference
Senior Solutions Architect, AI Inference

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior Solutions Architect, AI Inference
Senior Solutions Architect, AI Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior AI Infra Architect for Enterprise GPU Clusters
Senior AI Infra Architect for Enterprise GPU Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 288,000
Equity
Benefits