Low-Latency Inference Systems Engineer (Multi-GPU)

Digital Waffle

San Francisco (CA)

Hybrid

USD 180,000 - 280,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Digital Waffle in San Francisco is seeking a systems-focused engineer to own and optimize large-scale model serving. You will work on a fork of vLLM across an H100 fleet, focusing on latency and cost per token, with equity as part of the compensation.

The role involves kernel-level optimization, advanced batching strategies, and careful quantisation decisions to meet customer latency targets. A hybrid work setup is offered, and growth toward leading a larger inference team is expected.

Qualifications

  • Backgrounds that translate: HPC, graphics, embedded, compiler work, or performance engineering where latency was a product requirement.

Responsibilities

  • Kernel-level optimization using CUDA and Triton, custom ops, fused attention.
  • Implement continuous batching, KV-cache strategy, and speculative decoding tailored per workload.
  • Evaluate quantisation trade-offs where customers notice quality loss first.
  • Perform multi-GPU and multi-node serving with profiling to prove improvements.

Skills

Latency optimization
Performance engineering
Multi-GPU scaling

Tools

CUDA
Triton
custom ops

Job description

Digital Waffle in San Francisco is seeking a systems-focused engineer to own and optimize large-scale model serving. You will work on a fork of vLLM across an H100 fleet, focusing on latency and cost per token, with equity as part of the compensation.

The role involves kernel-level optimization, advanced batching strategies, and careful quantisation decisions to meet customer latency targets. A hybrid work setup is offered, and growth toward leading a larger inference team is expected.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Inference Software Engineer
Inference Software Engineer

Digital Waffle • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
Scale Low-Latency ML Inference Engineer (GPU/CUDA)
Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Performance Engineer: GPU Kernel & Inference Optimize
Performance Engineer: GPU Kernel & Inference Optimize

WORLD LABS • San Francisco (CA)

On-site
USD 200,000 - 300,000
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Machine Learning Engineer, LLM Inference Optimization in San Francisco
Machine Learning Engineer, LLM Inference Optimization in San Francisco

Energy Jobline ZR • San Francisco (CA)

On-site
USD 180,000 - 300,000
Low-Latency ML Inference Engineer (GPU/FPGA)
Low-Latency ML Inference Engineer (GPU/FPGA)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences