GPU Inference Performance Manager

Candidate

San Francisco (CA)

On-site

USD 220,000 - 270,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
US medical/dental/vision
Flexible PTO
Parental leave
Fertility stipend
401(k)

Job summary

Baseten is seeking an Engineering Manager to lead part of our Inference Performance team in San Francisco. You will mentor engineers, set direction for GPU optimization across the inference engine and runtime, and partner with cross-functional groups to ship high-performance AI workloads.

This hands-on leadership role requires deep technical depth in GPUs, experience hiring top talent, and the ability to translate performance wins into measurable outcomes such as lower latency and cheaper

Qualifications

  • Bachelor's, Master's, or Ph.D. in CS, Engineering, Mathematics, or related field.
  • Experience managing engineers, including hiring, mentoring, giving feedback and performance reviews.
  • Experience leading or supporting GPU optimization teams in training, inference or recommendation systems.
  • Strong technical depth in GPU workloads, with understanding of GPU architecture and tradeoffs.
  • Familiarity with ML libraries such as PyTorch, TensorRT or TensorRT-LLM.
  • Track record of driving roadmaps and shipping complex technical projects with a team.
  • Clear written and verbal communication to align stakeholders across teams.

Responsibilities

  • Lead, mentor and grow a team of inference performance engineers through regular 1:1s and performance reviews.
  • Hire top GPU and inference talent and build a collaborative team culture as the runtime team scales.
  • Own the technical roadmap and execution for runtime performance work.
  • Review designs, guide profiling and optimization, and reason from first principles about time and memory.
  • Drive the productionization of inference techniques such as quantization, KV-cache reuse, and scheduling.
  • Turn performance wins into measurable outcomes like tokens per GPU-hour, latency, and cost.

Skills

GPU optimization
People management
Hiring & mentoring
Performance reviews
PyTorch
TensorRT
CUDA/Triton

Education

Bachelor's/Master's/Ph.D. in CS/Engineering/Math
Advanced degree preferred

Tools

PyTorch
TensorRT
CUDA
Triton
TensorRT-LLM

Job description

Baseten is seeking an Engineering Manager to lead part of our Inference Performance team in San Francisco. You will mentor engineers, set direction for GPU optimization across the inference engine and runtime, and partner with cross-functional groups to ship high-performance AI workloads.

This hands-on leadership role requires deep technical depth in GPUs, experience hiring top talent, and the ability to translate performance wins into measurable outcomes such as lower latency and cheaper

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Hands-on Engineering Manager, GPU Inference
Hands-on Engineering Manager, GPU Inference

Baseten • United States

Remote
USD 190,000 - 260,000
Engineering Manager - Inference Performance
Engineering Manager - Inference Performance

Baseten • United States

Remote
USD 190,000 - 260,000
Engineering Manager - Inference Performance
Engineering Manager - Inference Performance

Baseten • United States

Remote
USD 190,000 - 260,000
Engineering Manager, GPU AI Inference at Scale
Engineering Manager, GPU AI Inference at Scale

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 431,250
Equity
Benefits
Engineering Manager, AI Inference & GPU Scaling
Engineering Manager, AI Inference & GPU Scaling

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 431,250
Equity
Benefits
Technical Program Manager, AI Inference & Model Performance
Technical Program Manager, AI Inference & Model Performance

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Equity
Medical, dental, vision insurance
Flexible PTO
+3
Technical Program Manager, AI Inference & Model Performance
Technical Program Manager, AI Inference & Model Performance

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Equity
Medical, dental, vision insurance
Flexible PTO
+3