Senior AI Inference Architect — Disaggregated Kubernetes

NVIDIA

United States

Remote

USD 184,000 - 357,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is hiring an architect to reshape AI inference with disaggregated systems and scalable reference architectures. You will guide partners, design end-to-end inference pipelines, and optimize performance across GPU clusters at scale.

You will mentor teams, drive strong customer engagements, and contribute to the deployment of cutting-edge NVIDIA Dynamo, Triton Inference Server, and TensorRT-LLM based solutions. Equity and comprehensive benefits are provided.

Qualifications

  • 6+ Years in Solutions Architecture driving customer engagements to deploy distributed systems and AI workloads.
  • 2+ years with AI workloads on Kubernetes.
  • Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization/serving.
  • Deep knowledge of disaggregated serving, multi-tier KV cache management, speculative decoding, quantization, and custom inference kernels.
  • Hands-on full-stack agent design with modern sandboxing, memory and retrieval systems with evaluation, skill design and governance.

Responsibilities

  • Work with inference partners to build and operate inference recipes and improve efficiency.
  • Accelerate inference pipelines using TensorRT-LLM, vLLM, and other backends to integrate with disaggregated inference.
  • Evangelize devops best-practices: Kubernetes clusters, compute fabrics, observability.
  • Provide mentorship and technical leadership to customers and internal teams during deployment of disaggregated inference systems.

Skills

Solutions Architecture
Customer engagements
AI workloads on Kubernetes
Full-stack agent design
Mentorship & leadership

Education

BS in CS/Engineering

Tools

NVIDIA Dynamo
Triton Inference Server
TensorRT-LLM
Kubernetes
Disaggregated inference

Job description

NVIDIA is hiring an architect to reshape AI inference with disaggregated systems and scalable reference architectures. You will guide partners, design end-to-end inference pipelines, and optimize performance across GPU clusters at scale.

You will mentor teams, drive strong customer engagements, and contribute to the deployment of cutting-edge NVIDIA Dynamo, Triton Inference Server, and TensorRT-LLM based solutions. Equity and comprehensive benefits are provided.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Architect — Scale & Disaggregated Serving
Senior AI Inference Architect — Scale & Disaggregated Serving

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Lead Architect, AI Inference & Disaggregated Serving
Lead Architect, AI Inference & Disaggregated Serving

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Lead Solutions Architect — AI Inference & Disaggregated Serving
Lead Solutions Architect — AI Inference & Disaggregated Serving

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Competitive salaries
Comprehensive benefits package
Equal opportunity employer
Senior AI Infra Architect: Kubernetes at Scale (Equity)
Senior AI Infra Architect: Kubernetes at Scale (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 224,000 - 357,000
Equity
Benefits
Senior Solutions Architect, AI Inference
Senior Solutions Architect, AI Inference

NVIDIA • United States

Remote
USD 184,000 - 357,000
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Senior Solutions Architect, AI Inference
Senior Solutions Architect, AI Inference

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior Architect, Scaled AI Inference
Senior Architect, Scaled AI Inference

NVIDIA • United States

On-site
USD 200,000 - 320,000
Benefits package
Senior Solutions Architect, AI Inference
Senior Solutions Architect, AI Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior Solutions Architect, Inference Service Providers
Senior Solutions Architect, Inference Service Providers

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 184,000 - 357,000
Competitive salaries
Comprehensive benefits package
Equal opportunity employer