Staff Engineer, Local Inference & Kernel Optimization

Sonder

New York (NY)

On-site

USD 250,000 - 300,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision coverage
Flexible PTO
Relocation support

Job summary

Sonder is seeking an inference and performance engineer to own the systems layer between our models and the hardware they run on—starting with Apple Silicon and macOS. You will make our models faster, smaller, and more power-efficient through kernel optimization, memory layouts, and quantization.

You will work closely with researchers to co-design inference systems, profiling from workloads to measured improvements, and shipping reliable local inference stacks across the Mac platform.

Qualifications

  • Deep experience in ML inference, GPU programming, or numerical computing with shipped improvements.
  • Strong C++ skills and hands-on kernel writing in Metal, CUDA, Triton, or similar environments.
  • Understanding GPU architectures and memory hierarchies, bandwidth, and synchronization impacts on kernels.
  • Experience optimizing matrix multiplication, attention, or similar heavy ops, with numerical validation.

Responsibilities

  • Own end-to-end performance: profile latency, throughput, memory, and energy across workloads.
  • Build kernels where it matters: optimize core ops with tiling and memory layouts to reduce traffic.
  • Make low-precision inference useful: evaluate quantization and mixed precision for performance and quality.
  • Improve the runtime: reduce overhead and enable reuse of computation across screen content changes.
  • Co-design for consumer hardware: align architecture decisions with Mac memory and compute.
  • Build performance infrastructure: reproducible benchmarks and checks across Apple Silicon generations.

Skills

ML inference
GPU programming
C++
Kernel optimization
Profiling
Quantization
Model optimization
System-level thinking

Tools

Metal
CUDA
Triton

Job description

Sonder is seeking an inference and performance engineer to own the systems layer between our models and the hardware they run on—starting with Apple Silicon and macOS. You will make our models faster, smaller, and more power-efficient through kernel optimization, memory layouts, and quantization.

You will work closely with researchers to co-design inference systems, profiling from workloads to measured improvements, and shipping reliable local inference stacks across the Mac platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Local Inference & Kernels
Member of Technical Staff, Local Inference & Kernels

Sonder • New York (NY)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
Flexible PTO
Relocation support
Silicon Kernel Performance Engineer for Power & Efficiency
Silicon Kernel Performance Engineer for Power & Efficiency

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 129,000 - 225,000
Medical & Dental
Retirement benefits
Relocation assistance
+1
Founding Inference Engineer, Apple Silicon Platform
Founding Inference Engineer, Apple Silicon Platform

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Senior ML Engineer, Foundation Models Inference — Cloud OS
Senior ML Engineer, Foundation Models Inference — Cloud OS

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock programs
Member of Technical Staff (Inference) - AI Infrastructure
Member of Technical Staff (Inference) - AI Infrastructure

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits
Senior ML Engineer, Foundation Model Inference (Cloud OS)
Senior ML Engineer, Foundation Model Inference (Cloud OS)

Apple Inc. • Seattle (WA)

On-site
USD 185,000 - 325,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
AI/ML Performance Engineer — Apple Silicon
AI/ML Performance Engineer — Apple Silicon

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Performance Engineer, Inference Engine - Flexible Hours
Performance Engineer, Inference Engine - Flexible Hours

Anthropic • San Francisco (CA), New York (NY)

On-site
USD 350,000 - 850,000