Senior Deep Learning Inference Engineer - Distributed Systems

NVIDIA AI

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA seeks a Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-latency AI services. You’ll lead architecture and performance work across Dynamo and open source frameworks, collaborating with research, software, systems, and hardware teams to make AI inference faster, more efficient, and easier to deploy.

Responsibilities include designing and maintaining Dynamo integrations with vLLM, SGLang, and TRTLLM; reducing

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent).
  • 3+ years building, profiling, and debugging performance-critical distributed or ML systems.
  • Strong programming skills in Python and/or Rust, C++.

Responsibilities

  • Design, build, and maintain Dynamo integrations for open source frameworks vLLM, SGLang, TRTLLM.
  • Partner with open source communities to land measurable gains in latency, throughput, reliability, and efficiency.
  • Showcase NVIDIA token/watt leadership by pushing the pareto frontier on public/private benchmarks
  • Find and remove bottlenecks across runtimes, kernels, networking, routing, and orchestration.
  • Develop inference optimizations for scheduling, disaggregation, KV caching, and autoscaling.

Skills

Python
Rust
C++

Education

BS/MS/PhD in CS/EE/related

Job description

NVIDIA seeks a Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-latency AI services. You’ll lead architecture and performance work across Dynamo and open source frameworks, collaborating with research, software, systems, and hardware teams to make AI inference faster, more efficient, and easier to deploy.

Responsibilities include designing and maintaining Dynamo integrations with vLLM, SGLang, and TRTLLM; reducing

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Deep Learning Inference Architect (Distributed)
Senior Deep Learning Inference Architect (Distributed)

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Inference Architect for Distributed DL Systems
Senior AI Inference Architect for Distributed DL Systems

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000
Equity
Competitive benefits
Senior DL Inference Architect (Distributed AI) - Equity
Senior DL Inference Architect (Distributed AI) - Equity

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Senior DL Inference Architect (Hybrid)
Senior DL Inference Architect (Hybrid)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000
Senior DL Inference Architect — Open-Source + Equity
Senior DL Inference Architect — Open-Source + Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity and benefits
Career growth
Senior System Software Engineer - AI Inference Platform
Senior System Software Engineer - AI Inference Platform

NVIDIA • Santa Clara (CA)

Hybrid
USD 272,000 - 431,250
Equity
Benefits
Hybrid work
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity and benefits
Career growth
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000