Senior LLM Efficiency Architect — Model-System Optimizer

NVIDIA

Santa Clara (CA)

Hybrid

USD 184,000 - 357,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NVIDIA is seeking a strong technical leader to drive a unified strategy for making LLMs more efficient from research through deployment. You will lead a multidisciplinary effort combining model ideas, systems expertise and hardware awareness to deliver scalable improvements while managing compute, memory, power and cost constraints.

The role requires hands-on leadership across model, software and hardware roadmaps, with a focus on measurable throughput and cost-per-token improvements.

Qualifications

  • Advanced degree or equivalent in CS/EE or related field.
  • Strong background in AI systems and model architectures.
  • Proven track record in performance optimization and HPC.

Responsibilities

  • Lead cross-layer efforts to improve LLM efficiency across architecture, training and inference.
  • Analyze workloads mapping to GPUs, memory, interconnects and distributed systems.
  • Establish measurement-driven roadmaps from research to production.
  • Collaborate with researchers, engineers, compilers and hardware architects.

Skills

AI systems
Model architecture
High-performance computing
Hardware awareness
Performance optimization

Education

MS or PhD degree or equivalent

Job description

NVIDIA is seeking a strong technical leader to drive a unified strategy for making LLMs more efficient from research through deployment. You will lead a multidisciplinary effort combining model ideas, systems expertise and hardware awareness to deliver scalable improvements while managing compute, memory, power and cost constraints.

The role requires hands-on leadership across model, software and hardware roadmaps, with a focus on measurable throughput and cost-per-token improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Efficiency Architect: Model Systems Co-Design
Senior LLM Efficiency Architect: Model Systems Co-Design

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
LLM Efficiency Architect: Cross-Layer Performance Leader
LLM Efficiency Architect: Cross-Layer Performance Leader

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Training Performance Architect
Senior LLM Training Performance Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 356,500
Equity
Benefits
Senior DL Performance Efficiency Architect
Senior DL Performance Efficiency Architect

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Lead Architect, High-Performance LLM Training
Lead Architect, High-Performance LLM Training

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Senior LLM Infra Engineer — AI Model Serving
Senior LLM Infra Engineer — AI Model Serving

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits package
Competitive salary