Senior Remote LLM Inference Optimization Lead

Dragonfly Digital Management, LLC (Dragonfly Capital)

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dragonfly Digital Management, LLC (Dragonfly Capital) seeks a Senior Inference Optimization Engineer to work remotely on GPU infrastructure optimization, LLM inference performance, and cost efficiency for a privacy-first consumer AI portfolio.

The role focuses on optimizing GPU throughput, evaluating inference engines, and applying cutting-edge techniques such as continuous batching and cache management to scale high-volume workloads.

Qualifications

  • 5+ years in performance optimization or HPC with GPU architecture experience.
  • Hands-on experience with production LLM inference engines at high volume.
  • Experience with LLM optimization techniques such as continuous batching and KV cache management.
  • Proficiency in Python, Rust, or Go; C++/CUDA is a strong plus.

Responsibilities

  • Optimize GPU infrastructure and improve throughput and cost per token for LLM inference workloads
  • Build benchmarking harnesses to identify optimal inference engines and strategies for various workloads
  • Evaluate and implement emerging inference optimization techniques and hardware solutions

Skills

Performance optimization
HPC
LLM inference
Continuous batching
KV cache management
Python
Rust
Go
C++/CUDA

Tools

Nsight Systems
PyTorch Profiler

Job description

Dragonfly Digital Management, LLC (Dragonfly Capital) seeks a Senior Inference Optimization Engineer to work remotely on GPU infrastructure optimization, LLM inference performance, and cost efficiency for a privacy-first consumer AI portfolio.

The role focuses on optimizing GPU throughput, evaluating inference engines, and applying cutting-edge techniques such as continuous batching and cache management to scale high-volume workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Remote Inference Optimization Engineer
Remote Inference Optimization Engineer

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior LLM Inference Architect — Edge, Data Center, Remote
Senior LLM Inference Architect — Edge, Data Center, Remote

Cerence AI • United States

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Engineer — End-to-End Optimizer
Senior LLM Inference Engineer — End-to-End Optimizer

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Equity
Benefits
Remote LLM Inference Optimization Architect
Remote LLM Inference Optimization Architect

Modular • United States

Hybrid
USD 198,000 - 286,000
Stock options
Health insurance
401k matching
+2
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1