Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA

United States

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is advancing the frontier of large language models and their practical deployment. We seek outstanding engineers to join our team to push the performance and efficiency of LLM inference across datacenter-scale workloads.

You will work on scalable systems, run workloads to optimize latency and throughput, and collaborate across multiple teams to define requirements for state-of-the-art AI solutions.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience).
  • 15+ years of experience in deep learning and deep learning systems design.
  • Proficiency in Python and C++ programming
  • Strong understanding of computer architecture, and GPU/parallel datacenter computing fundamentals.
  • Proven interest in analyzing, modeling, and tuning application performance.

Responsibilities

  • Research and Development: Explore and incorporate contemporary research on generative AI, agents, and inference systems into the NVIDIA LLM software stack.
  • Workload Analysis and Optimization: Profile and optimize agentic LLM workloads to reduce latency and increase throughput while maintaining fidelity.
  • System Design and Implementation: Design scalable systems to accelerate agentic workflows and handle datacenter-scale use cases.
  • Collaboration and Communication: Advise iterations of NVIDIA software, hardware, and systems by engaging with teams at NVIDIA and external partners and formalizing requirements.

Skills

Python
C++
Deep learning systems
GPU/parallel computing
Performance modeling

Education

BS/MS/PhD in Computer Science / Electrical Engineering / Computer Engineering

Tools

CUDA
OpenCL

Job description

NVIDIA is advancing the frontier of large language models and their practical deployment. We seek outstanding engineers to join our team to push the performance and efficiency of LLM inference across datacenter-scale workloads.

You will work on scalable systems, run workloads to optimize latency and throughput, and collaborate across multiple teams to define requirements for state-of-the-art AI solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Architect — Equity Eligible, Remote
Senior LLM Inference Architect — Equity Eligible, Remote

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Senior LLM Inference Architect
Senior LLM Inference Architect

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Architect — Edge, Data Center, Remote
Senior LLM Inference Architect — Edge, Data Center, Remote

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Manager, Large Language Model Inference
Manager, Large Language Model Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Competitive salary
Equity options
Comprehensive benefits
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1