LLM Inference Engineer: GPU-Optimized Serving

Energy Jobline ZR

San Francisco (CA)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GMI Cloud in San Francisco is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

GMI Cloud invites engineers to work on frontier research across quantization, speculative decoding, KV cache & memory management, and PD disaggregation, partnering with platform,

Qualifications

  • Experience optimizing large-scale LLM inference in production.
  • Proven track record reducing latency and boosting throughput.
  • Contributions to open-source inference projects.

Responsibilities

  • Drive frontier research in LLM inference optimization.
  • Develop optimization strategies for large-scale serving across model execution, runtime systems, and production inference platforms.

Skills

LLM inference systems
inference metrics
serving stacks
GPU-based inference
Speculative decoding
Quantization
PD disaggregation
KV cache & memory
Open-source contributions
Technical publishing

Education

PhD in related areas

Tools

SGLang
vLLM
TensorRT-LLM
NVIDIA Dynamo
Triton

Job description

GMI Cloud in San Francisco is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

GMI Cloud invites engineers to work on frontier research across quantization, speculative decoding, KV cache & memory management, and PD disaggregation, partnering with platform,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Machine Learning Engineer (LLM inference)
Machine Learning Engineer (LLM inference)

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000
LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in San Francisco
Machine Learning Engineer, LLM Inference Optimization in San Francisco

Energy Jobline ZR • San Francisco (CA)

On-site
USD 180,000 - 300,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000