Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent

San Francisco (CA)

On-site

USD 180,000 - 220,000

Full time

39 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Acceler8 Talent is seeking a Member of Technical Staff, Inference Performance to accelerate AI model inference across heterogeneous hardware. You will drive optimizations in latency, throughput, and memory usage, and push the serving stack with batching, caching, and quantization techniques.

The role involves developing CUDA/HIP/Triton kernels, deploying on accelerator clusters, and benchmarking models under real production workloads to improve efficiency and cost effectiveness.

Qualifications

  • Evidence of work in high-performance AI inference or model-serving systems.
  • Experience with GPU kernel development using CUDA, HIP, or Triton.
  • Knowledge of quantization, speculative decoding, batching, or KV-cache optimization.
  • Familiarity with PyTorch, vLLM, SGLang, TensorRT-LLM, or similar frameworks.
  • Experience with GPU, accelerator, HPC, or distributed-compute infrastructure.
  • Strong background in low-level performance profiling and systems optimization.
  • Ability to operate latency-sensitive systems in production.

Responsibilities

  • Profile inference latency, throughput, and memory usage.
  • Improve serving engines through batching, caching, quantization, and speculative decoding.
  • Develop and tune CUDA/HIP/Triton kernels.
  • Operate heterogeneous accelerator clusters.
  • Benchmark models and hardware under production workloads.
  • Build observability and reliability into the serving platform.
  • Debug performance across models, runtimes, kernels, networking, and hardware.

Skills

Inference performance
GPU kernel development
Quantization
PyTorch
TensorRT-LLM
CUDA
HIP
Triton
Model serving
Profiling
Low-level optimization
Distributed compute

Job description

Acceler8 Talent is seeking a Member of Technical Staff, Inference Performance to accelerate AI model inference across heterogeneous hardware. You will drive optimizations in latency, throughput, and memory usage, and push the serving stack with batching, caching, and quantization techniques.

The role involves developing CUDA/HIP/Triton kernels, deploying on accelerator clusters, and benchmarking models under real production workloads to improve efficiency and cost effectiveness.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits