LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE

Sonoma (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NLP PEOPLE is seeking world-class Machine Learning Engineers to develop leading inference optimization solutions. The role involves driving research and executing advanced LLM techniques to establish industry benchmarks.

Candidates should have hands-on experience with GPU optimization and knowledge of serving technologies. A collaborative spirit and the ability to publish results are essential. This position represents an exciting opportunity to push the boundaries in AI technology.

Qualifications

  • Experienced with LLM inference systems and modern GPU optimization.
  • Familiar with inference metrics and tradeoffs.
  • Experience with frameworks for large-scale serving.

Responsibilities

  • Drive research in LLM inference optimization across focus tracks.
  • Develop optimization strategies for large-scale model serving.
  • Engage with the open-source community.

Skills

LLM inference optimization
GPU performance optimization
Quantization
Speculative decoding
Collaboration with teams

Education

2+ years of relevant experience or PhD

Tools

SGLang
vLLM
TensorRT-LLM
NVIDIA Dynamo
Triton

Job description

NLP PEOPLE is seeking world-class Machine Learning Engineers to develop leading inference optimization solutions. The role involves driving research and executing advanced LLM techniques to establish industry benchmarks.

Candidates should have hands-on experience with GPU optimization and knowledge of serving technologies. A collaborative spirit and the ability to publish results are essential. This position represents an exciting opportunity to push the boundaries in AI technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2