Senior LLM Infra Engineer — AI Model Serving

Nvidia Corporation

Santa Clara (CA)

On-site

USD 184,000 - 288,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits package
Competitive salary

Job summary

NVIDIA in Santa Clara, CA, is hiring for a senior role focused on LLM serving infrastructure and platform API development. You will own the fine-tuning pipeline, evaluation harness, and observability layers while coordinating with ISVs and CSPs to deploy NVIDIA NIMs at scale.

Required: 7+ years in LLM serving infra, deep quantization knowledge, and strong collaboration across teams. A competitive salary and equity benefit package accompany this opportunity.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 7+ years of experience with LLM serving infrastructure (vLLM, TRT-LLM, or equivalent).
  • Deep understanding of model quantization and fine-tuning workflows.
  • Demonstrated history of developing robust platform APIs.
  • Strong collaborative skills and experience working in cross-functional teams.
  • Outstanding problem-solving skills and meticulous attention to detail.

Responsibilities

  • Compose and build the fine-tuning handoff pipeline, including LoRA adapter repackaging, re-quantization, and re-validation into NIM.
  • Develop the evaluation harness, ensuring models meet our high standards.
  • Implement the observability and attestation layer to produce auditable compliance artifacts.
  • Work in close partnership with ISVs and CSPs to roll out NVIDIA NIMs on a large scale.
  • Define and improve durable platform APIs, steering clear of one-off integrations.
  • Ensure flawless completion of projects through strict attention to detail and proven methodologies.

Skills

LLM serving infra
Model quantization
API development
Cross-functional collaboration
Problem solving

Education

Bachelor's degree in CS/engineering or equivalent

Tools

VLLM
TRT-LLM

Job description

NVIDIA in Santa Clara, CA, is hiring for a senior role focused on LLM serving infrastructure and platform API development. You will own the fine-tuning pipeline, evaluation harness, and observability layers while coordinating with ISVs and CSPs to deploy NVIDIA NIMs at scale.

Required: 7+ years in LLM serving infra, deep quantization knowledge, and strong collaboration across teams. A competitive salary and equity benefit package accompany this opportunity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior LLM Efficiency Architect: Model Systems Co-Design
Senior LLM Efficiency Architect: Model Systems Co-Design

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, LLM Inference & Deployment at Scale
Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Engineering Manager, LLM Inference & Deployment at Scale
Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 320,000 - 420,000
Equity
Benefits
Senior LLM Inference Architect — Edge, Data Center, Remote
Senior LLM Inference Architect — Edge, Data Center, Remote

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
LLM-Powered AI Infrastructure Architect (Equity)
LLM-Powered AI Infrastructure Architect (Equity)

NVIDIA • North Carolina

On-site
USD 152,000 - 288,000
Equity
Benefits
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior R&D Engineer (LLM / Agentic)
Senior R&D Engineer (LLM / Agentic)

SoftServe • Town of Poland (NY)

On-site
USD 120,000 - 150,000