Senior Architect, Scaled AI Inference & Orchestration

NVIDIA

California (MO)

On-site

USD 320,000 - 489,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a leader to drive the global strategy for scaled-out AI inferencing. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and reliability across enterprise and cloud environments.

You will define the technical roadmap for full lifecycle deployment, versioning, and automated scaling, collaborating with leadership to ensure peak efficiency on NVIDIA hardware.

Qualifications

  • 16+ years in technical roles with a recent long-term focus on AI infrastructure and large-scale inference orchestration.
  • 7-10+ years of leadership experience.
  • BS/MS or higher or equivalent experience in systems / software engineering, or related engineering fields.
  • Deep Technical Expertise: CUDA, GPUs, hardware acceleration, and cloud-native multi-tenant model serving.
  • Proven success delivering high-impact technically complex solutions with transparent resource utilization and performance insights.
  • Technical Leadership: align leadership and cross-functional teams toward architecture and execution.

Responsibilities

  • Architect distributed pipelines for high-throughput, low-latency inference across massive-scale workloads.
  • Lead hardware-software co-optimization and kernel-level performance tuning for production-grade model serving.
  • Influence open source projects (Dynamo, TensorRT-LLM, vLLM, Ray) to run on NVIDIA hardware.
  • Orchestrate model lifecycles including deployment, versioning, and intelligent scaling across cloud and data center environments.
  • Engage with customers, infrastructure providers, and partners to set industry standards for performance and availability.
  • Oversee end-to-end software and system lifecycle from ideation to evolution.

Skills

Leadership
GPU architecture
CUDA
Cloud-native
Performance tuning
Distributed systems

Education

BS/MS or higher in systems/software engineering

Tools

Kubernetes
TensorRT
Ray

Job description

NVIDIA is seeking a leader to drive the global strategy for scaled-out AI inferencing. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and reliability across enterprise and cloud environments.

You will define the technical roadmap for full lifecycle deployment, versioning, and automated scaling, collaborating with leadership to ensure peak efficiency on NVIDIA hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Architect, Scaled AI Inference & Systems — Equity
Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Senior AI Inference Architect — Scale & Disaggregated Serving
Senior AI Inference Architect — Scale & Disaggregated Serving

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior AI Inference Architect — Disaggregated Serving
Senior AI Inference Architect — Disaggregated Serving

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Lead Architect, AI Inference & Disaggregated Serving
Lead Architect, AI Inference & Disaggregated Serving

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Lead Solutions Architect — AI Inference & Disaggregated Serving
Lead Solutions Architect — AI Inference & Disaggregated Serving

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Competitive salaries
Comprehensive benefits package
Equal opportunity employer
Engineering Manager, GPU AI Inference at Scale
Engineering Manager, GPU AI Inference at Scale

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Architect: Kubernetes at Scale (Equity)
Senior AI Infra Architect: Kubernetes at Scale (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 224,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Architect for Enterprise ISVs
Senior AI Infrastructure Architect for Enterprise ISVs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Travel opportunities
Principal Infra Architect for AI-Driven Operations Equity
Principal Infra Architect for AI-Driven Operations Equity

NVIDIA AI • Redmond (WA)

On-site
USD 230,000 - 380,000
Competitive salary
Comprehensive benefits
Equity