Inference Software Engineer

Digital Waffle

San Francisco (CA)

Hybrid

USD 180,000 - 280,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Digital Waffle in San Francisco is seeking a systems-focused engineer to own and optimize large-scale model serving. You will work on a fork of vLLM across an H100 fleet, focusing on latency and cost per token, with equity as part of the compensation.

The role involves kernel-level optimization, advanced batching strategies, and careful quantisation decisions to meet customer latency targets. A hybrid work setup is offered, and growth toward leading a larger inference team is expected.

Qualifications

  • Backgrounds that translate: HPC, graphics, embedded, compiler work, or performance engineering where latency was a product requirement.

Responsibilities

  • Kernel-level optimization using CUDA and Triton, custom ops, fused attention.
  • Implement continuous batching, KV-cache strategy, and speculative decoding tailored per workload.
  • Evaluate quantisation trade-offs where customers notice quality loss first.
  • Perform multi-GPU and multi-node serving with profiling to prove improvements.

Skills

Latency optimization
Performance engineering
Multi-GPU scaling

Tools

CUDA
Triton
custom ops

Job description

Series A model-serving startup, ~35 people, $40M raised · San Francisco · Hybrid, 3 days · $180,000 - $280,000 + equity. They serve open-weight models for companies who can't send data to a third-party API. Their customers care about two numbers: p99 latency and cost per million tokens. This role owns both.

About the Role

The team runs a fork of vLLM across an H100 fleet. The last engineer to join cut p99 by 40% on the largest workload by reworking the scheduler. That's the standard.

Responsibilities
  • Kernel-level optimisation - CUDA and Triton, custom ops, fused attention
  • Continuous batching, KV-cache strategy and speculative decoding, tuned per workload rather than globally
  • Quantisation trade-offs where the customer notices quality loss before you do
  • Multi-GPU and multi-node serving, plus the profiling to prove any of it worked
Qualifications

Backgrounds that translate: HPC, graphics, embedded, compiler work, or performance engineering somewhere latency was a product requirement. Several of the strongest people in this space had never touched ML until it became a systems problem.

Required Skills

Four engineers on the inference team today, going to eight this year.

Pay range and compensation package

$180,000 - $280,000 + equity

Equal Opportunity Statement

We are committed to diversity and inclusivity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package
Senior Inference Platform Engineer - Data Center
Senior Inference Platform Engineer - Data Center

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff / Principal Machine Learning Engineer, Serving - USA
Staff / Principal Machine Learning Engineer, Serving - USA

Inworld • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options