LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE

Sonoma (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NLP PEOPLE is seeking world-class Machine Learning Engineers to develop leading inference optimization solutions. The role involves driving research and executing advanced LLM techniques to establish industry benchmarks.

Candidates should have hands-on experience with GPU optimization and knowledge of serving technologies. A collaborative spirit and the ability to publish results are essential. This position represents an exciting opportunity to push the boundaries in AI technology.

Qualifications

  • Experienced with LLM inference systems and modern GPU optimization.
  • Familiar with inference metrics and tradeoffs.
  • Experience with frameworks for large-scale serving.

Responsibilities

  • Drive research in LLM inference optimization across focus tracks.
  • Develop optimization strategies for large-scale model serving.
  • Engage with the open-source community.

Skills

LLM inference optimization
GPU performance optimization
Quantization
Speculative decoding
Collaboration with teams

Education

2+ years of relevant experience or PhD

Tools

SGLang
vLLM
TensorRT-LLM
NVIDIA Dynamo
Triton

Job description

NLP PEOPLE is seeking world-class Machine Learning Engineers to develop leading inference optimization solutions. The role involves driving research and executing advanced LLM techniques to establish industry benchmarks.

Candidates should have hands-on experience with GPU optimization and knowledge of serving technologies. A collaborative spirit and the ability to publish results are essential. This position represents an exciting opportunity to push the boundaries in AI technology.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Optimization Engineer - Frontier Performance
LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
LLM Inference & Systems Optimization Engineer
LLM Inference & Systems Optimization Engineer

Snowflake • Bellevue (KY)

On-site
USD 180,000 - 240,000
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Bonus and equity plan
Medical, dental, vision insurance
401(k) retirement plan
+1
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

On-site
USD 180,000 - 320,000
Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
LLM Inference Engineer (Mid, Senior, Staff)
LLM Inference Engineer (Mid, Senior, Staff)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2