Remote ML Inference Engineer - Serving & Performance

Yobitel Communications

United States

Remote

USD 160,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Yobitel Communications is seeking an engineer to own the model serving layer, optimizing latency and throughput for scale. You will manage vLLM, TensorRT-LLM, and Triton deployments, shaping the benchmarking discipline that keeps performance honest.

You will drive workload-specific serving strategies, quantization choices, and rigorous benchmarking to stay ahead in a fast-moving field. This role is remote-friendly and full-time, with a focus on high-performance inference at scale.

Qualifications

  • Shipped at least one production inference stack on specialized GPUs (e.g., H100s or A100s).
  • Familiar with vLLM, TensorRT-LLM, and Triton in production contexts.
  • Strong opinions about speculative decoding and model serving trade-offs.

Responsibilities

  • Own the serving stack selection per workload, including batching modes and caching strategies.
  • Lead quantization efforts (FP8, AWQ, GPTQ) and build eval harnesses to prove trade-offs.
  • Develop and run InferenceBench-style benchmarks to compare our serving against the field.

Skills

Serving-stack design
Quantization
Benchmarking

Job description

Yobitel Communications is seeking an engineer to own the model serving layer, optimizing latency and throughput for scale. You will manage vLLM, TensorRT-LLM, and Triton deployments, shaping the benchmarking discipline that keeps performance honest.

You will drive workload-specific serving strategies, quantization choices, and rigorous benchmarking to stay ahead in a fast-moving field. This role is remote-friendly and full-time, with a focus on high-performance inference at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote ML Inference Infrastructure Engineer
Remote ML Inference Infrastructure Engineer

United States Digital Space LLC • United States

Remote
USD 150,000 - 230,000
Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Remote ML Performance Engineer: Optimize Training Inference
Remote ML Performance Engineer: Optimize Training Inference

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Remote ML Systems Engineer – High-Performance Inference
Remote ML Systems Engineer – High-Performance Inference

Triwill Group • United States

Remote
USD 145,000 - 165,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer - Inference / Serving
Machine Learning Engineer - Inference / Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Competitive base salary
Meaningful equity
Annual bonus
+3
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Remote MLOps Engineer: Scale AI Inference & Serving
Remote MLOps Engineer: Scale AI Inference & Serving

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000