LLM Inference Optimization Engineer

GMI Cloud

Mountain View (CA)

On-site

USD 170,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GMI Cloud is building the leading inference optimization solution and the most advanced token platform for the global token market, hiring world‑class ML engineers to push the boundaries of LLM serving performance, cost efficiency, and reliability.

You will drive research, validation, and productionization of cutting-edge inference optimization techniques, turning ideas into measurable improvements over baselines and contributing back to the community.

Qualifications

  • Strong systems and infrastructure background with Linux, networking, observability and distributed systems.
  • Experience deploying production ML platforms and serving stacks.
  • Familiarity with LLM inference optimization and benchmark workflows.

Responsibilities

  • Design and operate a reliable, repeatable experiment framework for LLM inference across H200 and B200 fleets.
  • Build A/B testing infrastructure against real traffic.
  • Optimize and land optimized inference solutions for customized inference scenarios.
  • Build the foundation for agentic / automated inference optimization.
  • Drive reliability and fault recovery of advanced inference solution toward high SLA targets.
  • Build elastic GPU provisioning across hardware and spot machines to support production and experiments.
  • Diagnose and fix long-tail distributed inference performance bugs.
  • Collaborate with ML engineers to benchmark, test, and roll out optimizations.
  • Engage with open-source community (NVIDIA Dynamo, NCCL, vLLM, SGLang, TensorRT-LLM).

Skills

Linux / distributed systems
Python
Go / C++ / Rust
GPU infrastructure
Experimentation platforms
Hardware/Software debugging

Tools

NVIDIA Dynamo
vLLM
TensorRT-LLM
NCCL
InfiniBand
RoCE

Job description

GMI Cloud is building the leading inference optimization solution and the most advanced token platform for the global token market, hiring world‑class ML engineers to push the boundaries of LLM serving performance, cost efficiency, and reliability.

You will drive research, validation, and productionization of cutting-edge inference optimization techniques, turning ideas into measurable improvements over baselines and contributing back to the community.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Optimization Engineer
LLM Inference Optimization Engineer

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Edge-to-Cloud LLM Inference Engineer
Edge-to-Cloud LLM Inference Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1