Member of Technical Staff, Local Inference & Kernels

Sonder

New York (NY)

On-site

USD 250,000 - 300,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision coverage
Flexible PTO
Relocation support

Job summary

Sonder is seeking an inference and performance engineer to own the systems layer between our models and the hardware they run on—starting with Apple Silicon and macOS. You will make our models faster, smaller, and more power-efficient through kernel optimization, memory layouts, and quantization.

You will work closely with researchers to co-design inference systems, profiling from workloads to measured improvements, and shipping reliable local inference stacks across the Mac platform.

Qualifications

  • Deep experience in ML inference, GPU programming, or numerical computing with shipped improvements.
  • Strong C++ skills and hands-on kernel writing in Metal, CUDA, Triton, or similar environments.
  • Understanding GPU architectures and memory hierarchies, bandwidth, and synchronization impacts on kernels.
  • Experience optimizing matrix multiplication, attention, or similar heavy ops, with numerical validation.

Responsibilities

  • Own end-to-end performance: profile latency, throughput, memory, and energy across workloads.
  • Build kernels where it matters: optimize core ops with tiling and memory layouts to reduce traffic.
  • Make low-precision inference useful: evaluate quantization and mixed precision for performance and quality.
  • Improve the runtime: reduce overhead and enable reuse of computation across screen content changes.
  • Co-design for consumer hardware: align architecture decisions with Mac memory and compute.
  • Build performance infrastructure: reproducible benchmarks and checks across Apple Silicon generations.

Skills

ML inference
GPU programming
C++
Kernel optimization
Profiling
Quantization
Model optimization
System-level thinking

Tools

Metal
CUDA
Triton

Job description

About Sonder

Sonder is an applied AI lab building models that learn how people work.

Today, AI largely depends on people explaining what they are doing and remembering when to ask for help. We are building private models that live on your computer, understand work as it happens, and learn the patterns, preferences, and judgment behind how you operate.

Our first product, Twin, brings this intelligence to the Mac. It builds a memory of your work and uses that context to offer help at the right moment.

The Role

We are looking for an inference and performance engineer to own the systems layer between our models and the hardware they run on.

You will make our models faster, smaller, and more power-efficient, working from model execution and quantization down to memory layouts and custom GPU kernels. We are starting with Apple Silicon and macOS, using MLX and Metal.

Our models run throughout the working day, sharing memory and compute with the applications someone is using. Sustained power draw, memory bandwidth, and responsiveness under contention matter as much as peak throughput. A faster kernel matters when it makes the whole system better.

You will work closely with researchers to co-design models and inference systems around these constraints, with ownership from profiling an unfamiliar workload to shipping a measured improvement.

Why This Matters at Sonder

Local intelligence has to earn its place on someone's computer. Every reduction in memory use or sustained power draw makes room for more capable models and longer context while keeping the Mac responsive.

This work determines how much intelligence can stay private on the device, and how naturally Twin can fit into a person's working day.

What You'll Do
  • Own end-to-end performance. Profile latency, throughput, peak memory, and energy use across representative workloads. Separate time spent in vision encoding, prefill, and decoding, and identify whether the limiting factor is compute, memory bandwidth, or runtime overhead.
  • Build kernels where they matter. Optimize critical operations such as matrix multiplication, attention, and quantized linear layers. Use tiling, fusion, and appropriate tensor layouts to reduce memory traffic while managing register pressure, occupancy, and synchronization.
  • Make low-precision inference useful. Evaluate weight and KV-cache quantization, mixed precision, and dequantization costs against both performance and model quality. Work with researchers to understand which layers and workloads tolerate lower precision.
  • Improve the runtime. Reduce allocation and kernel-launch overhead, manage caches, and schedule latency-sensitive inference alongside background work. Find opportunities to reuse computation as screen content and context change, and validate when reuse remains correct.
  • Co-design for consumer hardware. Help researchers assess architecture and context-length choices against the memory and compute available on a Mac. Turn promising experiments into reliable implementations in our local inference stack.
  • Build performance infrastructure. Maintain reproducible benchmarks and numerical checks across supported Apple Silicon generations. Measure sustained behavior, thermal effects, and regressions alongside other applications, and use that evidence to decide what we build or contribute upstream.
What We're Looking For

You do not need prior Apple Silicon experience. We care about demonstrated ability to understand hardware and make neural networks run substantially better on it.

  • Deep experience in ML inference, GPU programming, or numerical computing, with concrete examples of improvements you have shipped.
  • Strong C++ skills and hands-on experience writing kernels in Metal, CUDA, Triton, or a comparable accelerator programming environment.
  • A working understanding of GPU architecture: memory hierarchies, bandwidth, SIMD execution, register pressure, occupancy, and synchronization. You can explain how these affect a kernel's performance.
  • Experience optimizing matrix multiplication, attention, or similarly demanding operations, including validating numerical correctness across shapes and precision formats.
  • An understanding of quantization and mixed precision, and the ability to measure their effects on memory, execution time, and model quality.
  • Strong profiling instincts. You can trace an application-level slowdown to the operation, memory access, or synchronization responsible, then verify the improvement in the full workload.
  • The judgment to move between model-level choices and low-level implementation, and to explain the tradeoffs clearly to researchers and engineers.
Particularly Exciting

We would be especially interested in depth in one or more of:

  • MLX internals, Metal Shading Language, or contributions to local inference engines such as llama.cpp.
  • Quantized matrix multiplication, fused attention, or efficient KV-cache management.
  • Vision-language models, incremental computation, or streaming inference.
  • Performance engineering for devices with tight battery, memory, or thermal budgets.

Evidence of technical depth matters more to us than matching every item.

Logistics
  • Location: This role is based in New York. We work together in person and are building the early team here.
  • Visa Sponsorship: We sponsor visas. We cannot guarantee success in every case, but if you are the right fit, we are committed to working through the process with you.
Compensation & Benefits
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $250,000–$300,000 USD, plus meaningful equity.
  • Benefits: Sonder offers comprehensive health, dental, and vision coverage, flexible PTO, and relocation support as needed.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Design
Member of Technical Staff, Design

Sonder • New York (NY)

On-site
USD 180,000 - 250,000
Health insurance
Relocation support
Flexible PTO
Member of Technical Staff, Product Engineer
Member of Technical Staff, Product Engineer

Sonder • New York (NY)

On-site
USD 180,000 - 250,000
Comprehensive health coverage
Dental coverage
Vision coverage
+2
Member of Technical Staff, Data and RL
Member of Technical Staff, Data and RL

Sonder • New York (NY)

On-site
USD 250,000 - 300,000
Health insurance
Dental insurance
Vision insurance
+2
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Staff Engineer, Local Inference & Kernel Optimization
Staff Engineer, Local Inference & Kernel Optimization

Sonder • New York (NY)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
Flexible PTO
Relocation support
Member of Technical Staff, Product Engineering
Member of Technical Staff, Product Engineering

SF Tensor • San Francisco (CA)

On-site
USD 225,000 - 275,000
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Member of Technical Staff, Product Engineering
Member of Technical Staff, Product Engineering

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 225,000 - 275,000
Relocation assistance
Equity
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1