Inference Infrastructure Engineer, Serving

Jobtailor

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Palo Alto is seeking an experienced ML inference engineer to build and optimize high-throughput inference services for large multimodal models. You will implement multi-GPU and multi-node parallelism, tune latency and throughput, and collaborate with researchers to push new architectures.

This on-site role emphasizes reliability, observability, and efficiency, with opportunities to shape production ML stacks and optimize GPU fleets at scale.

Qualifications

  • 3+ years of experience building low-latency inference serving systems for large models.
  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management).
  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang.
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel).
  • Strong systems programming skills; C++/CUDA a plus alongside Python.
  • Experience with autoscaling and load balancing for production ML services.
  • A track record of GPU cost optimization at scale.
  • Preferred qualifications (strong candidates may have some, not all): Experience serving multimodal models; contributions to open-source ML projects.

Responsibilities

  • Build low-latency, high-throughput inference serving systems for our large multimodal models
  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management
  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)
  • Build autoscaling and load balancing for production ML services
  • Establish standards for reliability, observability, and reproducibility across the inference stack
  • Collaborate with researchers to enable high-performance inference for novel architectures

Skills

Low-latency inference
GPU optimization
C++/CUDA
Python
Distributed systems
Model parallelism
Autoscaling

Tools

vLLM
TensorRT-LLM
Triton
SGLang

Job description


  • Build low-latency, high-throughput inference serving systems for our large multimodal models

  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management

  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory

  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)

  • Build autoscaling and load balancing for production ML services

  • Establish standards for reliability, observability, and reproducibility across the inference stack

  • Collaborate with researchers to enable high-performance inference for novel architectures


Requirements


  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models

  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)

  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang

  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)

  • Strong systems programming skills; C++/CUDA a plus alongside Python

  • Experience with autoscaling and load balancing for production ML services

  • A track record of GPU cost optimization at scale

  • Preferred qualifications (strong candidates may have some, not all): Experience serving multimodal (vision + language) models

  • Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior Inference Engineer
Senior Inference Engineer

Loft Labs, Inc. dba vCluster Labs • United States

On-site
USD 140,000 - 190,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits