Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA

Santa Clara (CA)

On-site

USD 320,000 - 489,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA is seeking a senior leader to shape the global strategy for scaled-out AI inference. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and reliability on NVIDIA hardware.

You will drive the technical roadmap across deployment, versioning, and automated scaling in enterprise and cloud environments, overseeing cross-functional teams and aligning with executive leadership.

Qualifications

  • 16+ years in technical roles with focus on AI infrastructure and large-scale inference orchestration
  • 7-10+ years of leadership experience
  • BS/MS or higher or equivalent in systems/software engineering or related fields
  • Deep technical expertise: GPU architecture, hardware acceleration, low-level performance tuning (CUDA, kernels) and multi-tenant cloud architectures
  • Proven success delivering complex, transparent, scalable distributed systems
  • Technical leadership: align cross-functional teams with senior leadership
  • Strong communication and teamwork in collaborating with customers and partners

Responsibilities

  • Architect distributed pipelines for high-throughput, low-latency AI inference at massive scale
  • Drive hardware-software co-optimization and performance tuning from kernel to driver
  • Guide open source and ecosystem projects (Dynamo, TensorRT-LLM, vLLM, SGLang, Ray) for NVIDIA hardware
  • Lead lifecycle management: automated deployment, versioning, and intelligent scaling across cloud and data center environments
  • Engage with customers, infrastructure providers, and partners to set industry standards for performance and availability
  • Oversee end-to-end software and system lifecycle from ideation to deployment and operations

Skills

Leadership
Communication
Strategy development
Cross-functional collaboration
Technical leadership
Architectural design

Education

BS/MS or higher in systems/software engineering

Tools

CUDA
Kernels
GPU architecture
Cloud-native architectures

Job description

NVIDIA is seeking a senior leader to shape the global strategy for scaled-out AI inference. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and reliability on NVIDIA hardware.

You will drive the technical roadmap across deployment, versioning, and automated scaling in enterprise and cloud environments, overseeing cross-functional teams and aligning with executive leadership.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Senior AI Inference Architect — Scale & Disaggregated Serving
Senior AI Inference Architect — Scale & Disaggregated Serving

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Lead Solutions Architect — AI Inference & Disaggregated Serving
Lead Solutions Architect — AI Inference & Disaggregated Serving

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Competitive salaries
Comprehensive benefits package
Equal opportunity employer
Lead Architect, AI Inference & Disaggregated Serving
Lead Architect, AI Inference & Disaggregated Serving

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior AI Infra Engineer — Scale AI Workloads & Equity
Senior AI Infra Engineer — Scale AI Workloads & Equity

Socket.dev • Santa Clara (UT)

On-site
USD 184,000 - 357,000
Equity
Health benefits
Senior AI Infra Architect: Kubernetes at Scale (Equity)
Senior AI Infra Architect: Kubernetes at Scale (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 224,000 - 357,000
Equity
Benefits
Engineering Manager, GPU AI Inference at Scale
Engineering Manager, GPU AI Inference at Scale

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Performance Engineer, Scale & AI — Equity
Senior GPU Performance Engineer, Scale & AI — Equity

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity and benefits
AI Inference Platform Lead – Equity Eligible
AI Inference Platform Lead – Equity Eligible

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package
Principal Infra Architect for AI-Driven Operations Equity
Principal Infra Architect for AI-Driven Operations Equity

NVIDIA AI • Redmond (WA)

On-site
USD 230,000 - 380,000
Competitive salary
Comprehensive benefits
Equity