Member of Technical Staff, Performance & GPU Kernels

ATBF Labs

San Francisco (CA)

Hybrid

USD 210,000 - 275,000

Full time

45 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health, dental, vision coverage
Equity

Job summary

ATBF Labs seeks a Member of Technical Staff to own performance at the metal, developing GPU kernels, memory movement, and numerics for low-precision safety. You will ship kernels that dramatically reduce latency and cost for production AI inference.

The role requires deep GPU architecture knowledge, CUDA/ROCm proficiency, and a track record of measurable speedups. You’ll collaborate with runtime and research teams to maximize hardware efficiency.

Qualifications

  • Experience writing and optimizing GPU kernels (CUDA, Triton, CUTLASS).
  • Profiling with Nsight/CUPTI to identify bottlenecks and optimize memory and compute.
  • Experience with mixed-precision or quantization schemes.
  • Familiarity with multi-GPU/multi-node communication (NVLink, InfiniBand, RoCE).
  • Strong understanding of GPU architecture, memory hierarchy, and parallel programming.
  • Experience with ML compilers and model inference stacks.

Responsibilities

  • Write and optimize GPU kernels (CUDA, Triton, CUTLASS) for attention, GEMMs, and quantized paths.
  • Profile with Nsight / CUPTI to find kernel- and memory-level bottlenecks, then close them.
  • Implement mixed-precision and quantization schemes (FP8, INT4, and beyond) with measured quality bounds.
  • Optimize communication for multi-GPU and multi-node serving over NVLink and RDMA (InfiniBand, RoCE).
  • Co-design model execution with the runtime and research teams for hardware efficiency.
  • Debug numerical instabilities that only appear at scale, on a fraction of requests.
  • Deep understanding of GPU architecture, parallel programming, and the memory hierarchy.
  • Proficiency in CUDA (or ROCm) and a GPU profiler (Nsight, nvprof, CUPTI).
  • Experience with performance-critical model execution and PyTorch internals.
  • A track record of measurable speedups, not just refactors.
  • Implement a fused attention kernel for a novel variant and beat the reference by a measurable margin.
  • Ship an FP8 path and run experiments to find the optimal speed/quality tradeoff.
  • Optimize RDMA communication patterns to remove a stall in multi-node decode.

Skills

CUDA
Triton
CUTLASS
Nsight
nvprof
CUPTI
PyTorch
GPU architecture
parallel programming
memory hierarchy
ROCm
torch.compile
XLA
TVM

Tools

Nsight
nvprof
CUPTI
torch.compile
Triton
XLA
TVM

Job description

Member of Technical Staff, Performance & GPU Kernels

You will take ownership of performance at the metal: GPU kernels, memory movement, and the numerics that make low precision safe. A single well-written kernel can change the economics of an entire model. You will find those wins and ship them.

About ATBF Labs

ATBF Labs builds the inference engine for production AI. Every token a model serves in production runs through an inference stack, and that stack decides the latency, the cost, and the reliability of the product sitting on top of it. We build ours from first principles: custom GPU kernels, a purpose-built runtime, and a distributed serving layer that holds its tail latency under real load. We are a small team with a high bar, shipping to production from day one.

What you'll do
Key responsibilities
  • Write and optimize GPU kernels (CUDA, Triton, CUTLASS) for attention, GEMMs, and quantized paths.
  • Profile with Nsight / CUPTI to find kernel- and memory-level bottlenecks, then close them.
  • Implement mixed-precision and quantization schemes (FP8, INT4, and beyond) with measured quality bounds.
  • Optimize communication for multi-GPU and multi-node serving over NVLink and RDMA (InfiniBand, RoCE).
  • Co-design model execution with the runtime and research teams for hardware efficiency.
  • Debug numerical instabilities that only appear at scale, on a fraction of requests.
  • Deep understanding of GPU architecture, parallel programming, and the memory hierarchy.
  • Proficiency in CUDA (or ROCm) and a GPU profiler (Nsight, nvprof, CUPTI).
  • Experience with performance-critical model execution and PyTorch internals.
  • A track record of measurable speedups, not just refactors.
  • Implement a fused attention kernel for a novel variant and beat the reference by a measurable margin.
  • Ship an FP8 path and run experiments to find the optimal speed/quality tradeoff.
  • Optimize RDMA communication patterns to remove a stall in multi-node decode.
Preferred qualifications
  • Experience with ML compilers (torch.compile, Triton, XLA, TVM).
  • Experience optimizing LLMs, VLMs, or video models for inference.
  • Familiarity with low-precision numerics and quantization-aware tradeoffs.
  • Contributions to open-source HPC or ML-systems projects.
Compensation

$210,000 – $275,000 + equity

Base salary plus meaningful equity. The range is a guideline; final numbers reflect experience, skills, and location. Full health, dental, and vision coverage included.

Why ATBF Labs
Solve hard problems

Inference is a systems problem from the kernel up. You will work on the parts that decide whether a model is usable in production: latency, throughput, and cost.

Own the whole stack

Small team, large surface area. You will have real ownership across kernels, runtime, and serving, and your work ships to customers, not a backlog.

Measure everything

We make decisions on numbers, not vibes. Every change is benchmarked, every regression is caught, and the survey point marks exactly where we are.

Learn from the best

Work alongside people who have built and operated inference at scale, and who care more about a clean result than a clever one.

ATBF Labs is an equal-opportunity employer. We celebrate diversity and are committed to an inclusive environment for everyone who builds with us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference Engine
Member of Technical Staff, Inference Engine

ATBF Labs • San Francisco (CA), Northern (KY)

On-site
USD 200,000 - 260,000
Health insurance
Member of Technical Staff, Research
Member of Technical Staff, Research

ATBF Labs • San Francisco (CA)

Hybrid
USD 215,000 - 285,000
Equity
Health benefits
Chief of Staff
Chief of Staff

ATBF Labs • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Engineering Manager - Inference Performance
Engineering Manager - Inference Performance

Xapply • San Francisco (CA)

On-site
USD 190,000 - 230,000
Equity
Health insurance (US)
Flexible PTO
+4
Engineering Manager - Inference Performance
Engineering Manager - Inference Performance

Candidate • San Francisco (CA)

On-site
USD 220,000 - 270,000
Equity
US medical/dental/vision
Flexible PTO
+3
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1