Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA

California (MO)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking outstanding engineers to help shape the future of LLM inference across agentic and reasoning use cases. We push the boundaries of what's possible with large language models and design scalable, datacenter-scale systems.

You will research and develop inference algorithms, profile workloads to reduce latency and increase throughput, and collaborate with diverse NVIDIA teams and partners to deliver efficient, cutting-edge solutions.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience).
  • 15+ years of experience in deep learning and deep learning systems design.
  • Proficiency in Python and C++ programming.
  • Strong understanding of computer architecture and GPU/parallel datacenter computing fundamentals.
  • Proven interest in analyzing, modeling, and tuning application performance.

Responsibilities

  • Research and Development: Explore and incorporate contemporary research on generative AI, agents, and inference systems into the NVIDIA LLM software stack.
  • Workload Analysis and Optimization: Profile and optimize agentic LLM workloads to reduce latency and increase throughput while maintaining fidelity.
  • System Design and Implementation: Design scalable systems to accelerate agentic workflows and handle datacenter-scale use cases.
  • Collaboration and Communication: Advise iterations of software, hardware, and systems by engaging with teams at NVIDIA and external partners.

Skills

Python
C++
Deep learning

Education

BS/MS/PhD in Computer Science, Electrical Engineering, Computer Engineering

Tools

CUDA
OpenCL

Job description

NVIDIA is seeking outstanding engineers to help shape the future of LLM inference across agentic and reasoning use cases. We push the boundaries of what's possible with large language models and design scalable, datacenter-scale systems.

You will research and develop inference algorithms, profile workloads to reduce latency and increase throughput, and collaborate with diverse NVIDIA teams and partners to deliver efficient, cutting-edge solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Architect — Equity Eligible, Remote
Senior LLM Inference Architect — Equity Eligible, Remote

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Senior LLM Inference Architect
Senior LLM Inference Architect

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Senior LLM Inference Architect — Performance & Benchmarking
Senior LLM Inference Architect — Performance & Benchmarking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior GPU AI Platform Engineer — Edge Inference (Equity)
Senior GPU AI Platform Engineer — Edge Inference (Equity)

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits