Member of Technical Staff, Inference Engine

ATBF Labs Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 200,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance

Job summary

ATBF Labs Inc. is building the core of production AI inference engines. You will own the runtime for batching, memory management, and low-latency serving across model families.

You will ship to production from day one, owning kernels, runtime, and the serving stack, with ownership across the stack and measurable impact on latency, throughput, and cost.

Qualifications

  • 5+ years building performance-critical systems in C++, Rust, CUDA, or similar.
  • Strong fundamentals: memory, concurrency, scheduling, and cache misses.
  • Hands-on experience serving deep-learning models in production.
  • Comfort reading flamegraphs and kernel traces to drive measurable wins.

Responsibilities

  • Design and build the inference runtime: continuous batching, paged attention/KV-cache management, and request scheduling.
  • Drive down tail latency (p99) and push up tokens-per-second across LLM, VLM, and multimodal workloads.
  • Implement and tune speculative decoding, prefix caching, and structured-output decoding paths.
  • Build the serving layer: load balancing, autoscaling, and graceful degradation under real production traffic.
  • Profile end to end, find the bottleneck, and fix it, whether in kernel, scheduler, or network.
  • Write benchmarking and regression harnesses for release quality control.

Skills

Performance systems
C/C++/Rust
Production DL systems
Flamegraphs

Tools

vLLM
TensorRT-LLM
SGLang
TGI

Job description

Member of Technical Staff, Inference Engine

You will build the core of our inference engine: the runtime that takes a set of weights and serves them as a low-latency, high-throughput endpoint. This is the layer where scheduling, batching, memory, and the model meet. Your work decides the latency every user feels and the cost of every token we serve.


About ATBF Labs

ATBF Labs builds the inference engine for production AI. Every token a model serves in production runs through an inference stack, and that stack decides the latency, the cost, and the reliability of the product sitting on top of it. We build ours from first principles: custom GPU kernels, a purpose-built runtime, and a distributed serving layer that holds its tail latency under real load. We are a small team with a high bar, shipping to production from day one.


What you'll do

Key responsibilities


  • Design and build the inference runtime: continuous batching, paged attention/KV-cache management, and request scheduling.

  • Drive down tail latency (p99) and push up tokens-per-second across LLM, VLM, and multimodal workloads.

  • Implement and tune speculative decoding, prefix caching, and structured-output decoding paths.

  • Build the serving layer: load balancing, autoscaling, and graceful degradation under real production traffic.

  • Profile end to end, find the bottleneck, and fix it, whether it lives in a kernel, the scheduler, or the network.

  • Write the benchmarking and regression harness that keeps every release honest.



  • 5+ years building performance-critical systems in C++, Rust, CUDA, or similar.

  • Strong systems fundamentals: memory, concurrency, scheduling, and the cost of a cache miss.

  • Hands-on experience serving deep-learning models in production, or building the systems that do.

  • Comfort reading a flamegraph and a kernel trace, and turning both into a measurable win.


Preferred qualifications


  • Experience with an inference framework (vLLM, TensorRT-LLM, SGLang, TGI) or building one.

  • Familiarity with the transformer serving path: attention, KV-cache, batching, quantization.

  • Experience with multi-GPU / multi-node serving and the collectives that make it work.

  • Open-source contributions to ML systems or runtimes.



  • Cut p99 latency for a 70B model by reworking the batching scheduler under bursty load.

  • Add a speculative-decoding path and measure the speed/quality tradeoff across draft models.

  • Build a KV-cache eviction policy that holds throughput as context lengths grow.


Compensation $200,000 – $260,000 + equity


Base salary plus meaningful equity. The range is a guideline; final numbers reflect experience, skills, and location. Full health, dental, and vision coverage included.


Why ATBF Labs

Solve hard problems

Inference is a systems problem from the kernel up. You will work on the parts that decide whether a model is usable in production: latency, throughput, and cost.


Own the whole stack

Small team, large surface area. You will have real ownership across kernels, runtime, and serving, and your work ships to customers, not a backlog.


Measure everything

We make decisions on numbers, not vibes. Every change is benchmarked, every regression is caught, and the survey point marks exactly where we are.


Learn from the best

Work alongside people who have built and operated inference at scale, and who care more about a clean result than a clever one.


ATBF Labs is an equal‑opportunity employer. We celebrate diversity and are committed to an inclusive environment for everyone who builds with us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Software Engineer, Model Performance Tooling
Software Engineer, Model Performance Tooling

BaseTen • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Equity
Health plans for dependents
Flexible PTO (Winter Break)
+3
Founding Engineer - ML Performance
Founding Engineer - ML Performance

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Forward Deployed Engineer (Training)
Forward Deployed Engineer (Training)

Baseten • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 260,000
Equity
Medical, dental, vision coverage for +
Flexible PTO including Winter Break
+4
Forward Deployed Engineer (Training)
Forward Deployed Engineer (Training)

The Consensus • New York (NY)

On-site
USD 120,000 - 170,000
Competitive compensation
Medical, dental, vision insurance
Flexible PTO