Infrastructure Engineer, LLM Inference Optimization

GMI Cloud

Mountain View (CA)

On-site

USD 170,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GMI Cloud is building the leading inference optimization solution and the most advanced token platform for the global token market, hiring world‑class ML engineers to push the boundaries of LLM serving performance, cost efficiency, and reliability.

You will drive research, validation, and productionization of cutting-edge inference optimization techniques, turning ideas into measurable improvements over baselines and contributing back to the community.

Qualifications

  • Strong systems and infrastructure background with Linux, networking, observability and distributed systems.
  • Experience deploying production ML platforms and serving stacks.
  • Familiarity with LLM inference optimization and benchmark workflows.

Responsibilities

  • Design and operate a reliable, repeatable experiment framework for LLM inference across H200 and B200 fleets.
  • Build A/B testing infrastructure against real traffic.
  • Optimize and land optimized inference solutions for customized inference scenarios.
  • Build the foundation for agentic / automated inference optimization.
  • Drive reliability and fault recovery of advanced inference solution toward high SLA targets.
  • Build elastic GPU provisioning across hardware and spot machines to support production and experiments.
  • Diagnose and fix long-tail distributed inference performance bugs.
  • Collaborate with ML engineers to benchmark, test, and roll out optimizations.
  • Engage with open-source community (NVIDIA Dynamo, NCCL, vLLM, SGLang, TensorRT-LLM).

Skills

Linux / distributed systems
Python
Go / C++ / Rust
GPU infrastructure
Experimentation platforms
Hardware/Software debugging

Tools

NVIDIA Dynamo
vLLM
TensorRT-LLM
NCCL
InfiniBand
RoCE

Job description

GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and one of only six cloud providers worldwide to earn NVIDIA's prestigious Reference Platform Cloud Partner designation. We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments. We empower AI startups and enterprises to "build AI without limits," providing everything they need to prototype, train, and deploy AI models quickly and reliably.

About this role

GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world-class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

This role is for engineers who want to live at the frontier of LLM inference systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage over top open-source baselines (vLLM, SGLang, and so on). Our charter is not just to adopt what's published — it is to define the recipes, ship the optimizations, and contribute back to the community that the rest of the industry follows.

Our goal is to make GMI the company that leads the industry in how fast we discover, evaluate, combine, and operationalize the best optimization strategies for real customer workloads. That means not only adopting the latest advances, but also defining best practices, developing our own optimization methodologies, and building the internal framework that keeps GMI ahead of the curve.

You will focus on B200-first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache and memory management, prefill/decode disaggregation, and system-level inference optimization. You will work closely with platform and infrastructure teams to transform cutting-edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability.

Key Responsibilities

  • Design and operate a reliable, repeatable experiment framework for LLM inference (vLLM / SGLang / TensorRT-LLM / NVIDIA Dynamo and in-house solution) across H200 and B200 fleets.
  • Build A/B testing infrastructure against real traffic.
  • Optimize and land optimized inference solution for customized inference scenarios.
  • Build the foundation for agentic / automated inference optimization.
  • Drive reliability and fault recovery of advance inference solution on top open-source models toward the SLA 99.99% target.
  • Build elastic GPU provisioning across H200, B200, and spot machines that powers both production traffic and experimentation.
  • Diagnose and fix the long tail of distributed-inference perf bugs: NCCL stalls, network contention, kernel-level latency.
  • Collaborate with ML engineers to make sure every optimization idea can be cleanly benchmarked, A/B tested, and rolled out.
  • Engage with and contribute to the open-source community (NVIDIA Dynamo, NCCL, vLLM, SGLang, TensorRT-LLM, AA benchmarks).

Required Skills

  • Strong systems / infrastructure engineering background with deep familiarity with Linux, networking, observability, and distributed systems.
  • Proficient in Python and at least one of Go / C++ / Rust.
  • Hands-on experience operating GPU clusters or building infrastructure for ML training or inference in production.
  • Hands on experience of experiment system for ML training or optimization.
  • For GPU & Network: hands-on experience with NCCL, InfiniBand or RoCE, GPUDirect RDMA, multi-node distributed training/inference debugging.
  • Comfortable owning end-to-end reliability — from hardware health and network topology up through the runtime and the application.
  • Strong debugging instincts: able to correlate symptoms across hardware counters, network telemetry, and application logs.
  • Clear communication; comfortable working with both ML engineers and hardware/network specialists.

Preferred Qualifications

  • 5+ years of experience of ML infra or building production ML platforms.
  • Experience benchmarking or productionizing modern serving stacks: vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, Triton.
  • Experience with H100 / H200 and especially B200 / Blackwell deployments.
  • Familiarity with PD disaggregation, KV cache offload, MoE serving topologies, or speculative decoding pipelines from an infrastructure standpoint.
  • Track record of contributing to open-source infra or benchmarking projects.
  • Experience publishing technical blogs or case studies from production infrastructure work.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site