Get more replies from employers
Send a job-specific resume in minutes.
Digital Waffle in San Francisco is seeking a systems-focused engineer to own and optimize large-scale model serving. You will work on a fork of vLLM across an H100 fleet, focusing on latency and cost per token, with equity as part of the compensation.
The role involves kernel-level optimization, advanced batching strategies, and careful quantisation decisions to meet customer latency targets. A hybrid work setup is offered, and growth toward leading a larger inference team is expected.
Series A model-serving startup, ~35 people, $40M raised · San Francisco · Hybrid, 3 days · $180,000 - $280,000 + equity. They serve open-weight models for companies who can't send data to a third-party API. Their customers care about two numbers: p99 latency and cost per million tokens. This role owns both.
The team runs a fork of vLLM across an H100 fleet. The last engineer to join cut p99 by 40% on the largest workload by reworking the scheduler. That's the standard.
Backgrounds that translate: HPC, graphics, embedded, compiler work, or performance engineering somewhere latency was a product requirement. Several of the strongest people in this space had never touched ML until it became a systems problem.
Four engineers on the inference team today, going to eight this year.
$180,000 - $280,000 + equity
We are committed to diversity and inclusivity.