LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

GMI Cloud, Inc is hiring world-class Machine Learning Engineers to advance LLM inference optimization leveraging GPUs and cutting-edge techniques. You will drive research, validation, and productionization of optimization strategies across NV platforms, with a focus on speed, efficiency, and scalability.

You will collaborate with platform and infrastructure teams, contribute to open-source projects, and help define recipes and benchmarks for industry-leading inference performance.

Qualifications

  • Must have hands-on experience with LLM inference systems and performance optimization.
  • Familiarity with inference metrics (TTFT, ITL, throughput, tail latency) and cost tradeoffs.
  • Experience with modern serving stacks (SGLang, vLLM, TensorRT-LLM, NVIDIA Dynamo, Triton).

Responsibilities

  • Drive frontier research and engineering in LLM inference optimization across specified tracks.
  • Develop optimization strategies for large-scale LLM serving with B200 primary target and H200 ongoing.
  • Collaborate with platform and infra teams to translate ideas into measurable performance gains.
  • Publish results and engage with open-source communities (vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo).
  • Build scalable optimization frameworks and benchmarking infrastructure.

Skills

LLM inference systems
Performance optimization
TTFT/ITL metrics
GPU-based inference
Serving stacks (SGLang, vLLM, TensorRT
Experimentation & benchmarks
Claude Code advanced usage
Technical communication

Education

PhD in related areas

Tools

SGLang
vLLM
TensorRT-LLM
NVIDIA Dynamo
Triton

Job description

GMI Cloud, Inc is hiring world-class Machine Learning Engineers to advance LLM inference optimization leveraging GPUs and cutting-edge techniques. You will drive research, validation, and productionization of optimization strategies across NV platforms, with a focus on speed, efficiency, and scalability.

You will collaborate with platform and infrastructure teams, contribute to open-source projects, and help define recipes and benchmarks for industry-leading inference performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity
Benefits
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
Remote Inference Optimization Engineer
Remote Inference Optimization Engineer

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

On-site
USD 180,000 - 320,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
Remote LLM Inference Optimization Architect
Remote LLM Inference Optimization Architect

Modular • United States

Hybrid
USD 198,000 - 286,000
Stock options
Health insurance
401k matching
+2
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
LLM Inference & Systems Optimization Engineer
LLM Inference & Systems Optimization Engineer

Snowflake • Bellevue (KY)

On-site
USD 180,000 - 240,000