Member of Technical Staff, Research

ATBF Labs

San Francisco (CA)

Hybrid

USD 215,000 - 285,000

Full time

38 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Health benefits

Job summary

ATBF Labs builds the inference engine for production AI. Every token a model serves in production runs through an inference stack emphasizing latency, throughput, cost, and quality. You will collaborate with the engine team to turn prototypes into production-grade systems, while benchmarks guide every decision.

The role emphasizes research with practical impact, including KV-cache compression, speculative decoding, and real-world deployment considerations.

Qualifications

  • PhD in CS, ML, or related field or equivalent research experience.
  • First-author publications at peer-reviewed venues.
  • Hedge hands-on experience with LLM/VLM inference internals.
  • Experience shipping research into production systems.

Responsibilities

  • Conduct research on efficient inference: speculative and parallel decoding, quantization, sparsity, and KV-cache compression.
  • Design, implement, and evaluate new methods against rigorous, reproducible benchmarks.
  • Partner with the engine team to transition prototypes into production-grade systems.
  • Analyze empirical results, find bottlenecks, and iterate quickly to improve model quality and speed.
  • Track emerging work and bring high-impact ideas into our roadmap.

Skills

ML research
Empirical methods
Python
C++/CUDA
Experiment design
Communication with engineers/research
Fluency with ML frameworks

Education

PhD in CS/ML or equivalent

Tools

PyTorch
JAX

Job description

You will push the boundary of efficient inference and carry the result all the way into production. We care about research that moves a real metric: latency, throughput, cost, or quality held constant while one of the others improves. You will collaborate closely with the engine team to turn a prototype into something serving live traffic.

About ATBF Labs

ATBF Labs builds the inference engine for production AI. Every token a model serves in production runs through an inference stack, and that stack decides the latency, the cost, and the reliability of the product sitting on top of it. We build ours from first principles: custom GPU kernels, a purpose-built runtime, and a distributed serving layer that holds its tail latency under real load. We are a small team with a high bar, shipping to production from day one.

What you’ll do
Key responsibilities
  • Conduct research on efficient inference: speculative and parallel decoding, quantization, sparsity, and KV-cache compression.
  • Design, implement, and evaluate new methods against rigorous, reproducible benchmarks.
  • Partner with the engine team to transition prototypes into production-grade systems.
  • Analyze empirical results, find the bottleneck, and iterate quickly to improve model quality and speed.
  • Track emerging work and bring the high-impact ideas into our roadmap.
Requirements
  • Research background in ML, systems, or a quantitative field, with a bias for empirical work.
  • Strong coding ability in Python and a systems language (C++/CUDA a plus).
  • Experience designing experiments and communicating results to engineers and researchers alike.
  • Fluency with at least one ML framework (PyTorch, JAX).
Preferred qualifications
  • PhD in CS, ML, or a related field, or equivalent research experience.
  • First-author publications at peer-reviewed venues (NeurIPS, ICML, MLSys, ICLR).
  • Hands-on work with LLM/VLM inference internals.
  • A history of shipping research into a real system, not just a paper.
  • Design a draft-model strategy for speculative decoding and prove the win on production traces.
  • Develop a KV-cache compression scheme and characterize the quality curve.
  • Build an evaluation harness that finds the optimal serving config for a class of models.

Compensation $215,000 – $285,000 + equity

Base salary plus meaningful equity. The range is a guideline; final numbers reflect experience, skills, and location. Full health, dental, and vision coverage included.

Why ATBF Labs
Solve hard problems

Inference is a systems problem from the kernel up. You will work on the parts that decide whether a model is usable in production: latency, throughput, and cost.

Own the whole stack

Small team, large surface area. You will have real ownership across kernels, runtime, and serving, and your work ships to customers, not a backlog.

Measure everything

We make decisions on numbers, not vibes. Every change is benchmarked, every regression is caught, and the survey point marks exactly where we are.

Learn from the best

Work alongside people who have built and operated inference at scale, and who care more about a clean result than a clever one.

ATBF Labs is an equal-opportunity employer. We celebrate diversity and are committed to an inclusive environment for everyone who builds with us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference Engine
Member of Technical Staff, Inference Engine

ATBF Labs • San Francisco (CA), Northern (KY)

On-site
USD 200,000 - 260,000
Health insurance
Chief of Staff
Chief of Staff

ATBF Labs • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
Member of Technical Staff, Performance & GPU Kernels
Member of Technical Staff, Performance & GPU Kernels

ATBF Labs • San Francisco (CA)

Hybrid
USD 210,000 - 275,000
Health, dental, vision coverage
Equity
Software Engineer- Inference Performance Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 25 minutes ago
Software Engineer- Inference Performance Baseten · New York City, NY Full-time · Hybrid $180,000–360,000 25 minutes ago

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
Software Engineer- Inference Performance
Software Engineer- Inference Performance

Baseten • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Meaningful equity
Medical, dental, vision insurance (US)
+4
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

On-site
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff ML Engineer - Efficient Production Inference
Staff ML Engineer - Efficient Production Inference

ATBF Labs • San Francisco (CA)

Hybrid
USD 215,000 - 285,000
Equity
Health benefits
Software Engineer- Inference Performance
Software Engineer- Inference Performance

Xapply • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive compensation and equity
Medical, dental, vision insurance (US)
Flexible PTO including Winter Break
+4
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1