LLM Inference Optimization Engineer

GMI Cloud

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market. We are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

This role focuses on frontier research, validation, and productionization of advanced inference optimization techniques, with close collaboration across platform and infrastructure teams to deliver

Qualifications

  • Experience with LLM inference systems and GPU performance optimization.
  • Understanding of inference metrics (throughput, latency, goodput) and tradeoffs.
  • Familiarity with modern serving stacks (SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, Triton).

Responsibilities

  • Drive frontier research and engineering in LLM inference optimization across focus tracks while contributing to the full optimization stack.
  • Develop next-generation optimization strategies for large-scale LLM serving across model execution and runtime systems.
  • Advance state-of-the-art techniques in quantization, speculative decoding, KV cache/memory management, and PD disaggregation.
  • Collaborate with platform, infrastructure, and product teams to translate ideas into measurable gains in latency, throughput, and cost efficiency.

Skills

LLM inference systems
GPU performance optimization
Inference metrics
Serving stacks
Experimentation

Education

PhD in related areas

Tools

SGLang
vLLM
TensorRT-LLM
NVIDIA Dynamo
Triton
CUDA

Job description

GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market. We are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

This role focuses on frontier research, validation, and productionization of advanced inference optimization techniques, with close collaboration across platform and infrastructure teams to deliver

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Edge-to-Cloud LLM Inference Engineer
Edge-to-Cloud LLM Inference Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1