Inference Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 180,000 - 220,000

Full time

27 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Acceler8 Talent is seeking a Member of Technical Staff, Inference Performance to accelerate AI model inference across heterogeneous hardware. You will drive optimizations in latency, throughput, and memory usage, and push the serving stack with batching, caching, and quantization techniques.

The role involves developing CUDA/HIP/Triton kernels, deploying on accelerator clusters, and benchmarking models under real production workloads to improve efficiency and cost effectiveness.

Qualifications

  • Evidence of work in high-performance AI inference or model-serving systems.
  • Experience with GPU kernel development using CUDA, HIP, or Triton.
  • Knowledge of quantization, speculative decoding, batching, or KV-cache optimization.
  • Familiarity with PyTorch, vLLM, SGLang, TensorRT-LLM, or similar frameworks.
  • Experience with GPU, accelerator, HPC, or distributed-compute infrastructure.
  • Strong background in low-level performance profiling and systems optimization.
  • Ability to operate latency-sensitive systems in production.

Responsibilities

  • Profile inference latency, throughput, and memory usage.
  • Improve serving engines through batching, caching, quantization, and speculative decoding.
  • Develop and tune CUDA/HIP/Triton kernels.
  • Operate heterogeneous accelerator clusters.
  • Benchmark models and hardware under production workloads.
  • Build observability and reliability into the serving platform.
  • Debug performance across models, runtimes, kernels, networking, and hardware.

Skills

Inference performance
GPU kernel development
Quantization
PyTorch
TensorRT-LLM
CUDA
HIP
Triton
Model serving
Profiling
Low-level optimization
Distributed compute

Job description

Member of Technical Staff, Inference Performance

200k base + equity

I’m working with a AI infrastructure startup building a high-performance inference cloud for open models.

The team optimizes the full path from model and serving engine through kernels, accelerators, and production infrastructure. Its platform is already processing trillions of tokens per month, and the company is expanding due to customer demand growing faster than its current capacity.

This role will focus on making inference faster, more reliable, and more cost-efficient across different models and hardware architectures.

You’ll work on problems such as:

  • Profiling and optimizing inference latency, throughput, and memory usage
  • Improving serving engines through batching, caching, quantization, and speculative decoding
  • Developing and tuning CUDA, HIP, or Triton kernels
  • Operating heterogeneous accelerator clusters
  • Benchmarking models and hardware under realistic production workloads
  • Building observability and reliability into the serving platform
  • Debugging performance across models, runtimes, kernels, networking, and hardware.

Looking for engineers who have strong evidence in one or more of:

  • High-performance AI inference or model-serving systems
  • GPU kernel development using CUDA, HIP, or Triton
  • Quantization, speculative decoding, batching, or KV-cache optimization
  • PyTorch, vLLM, SGLang, TensorRT-LLM, or similar frameworks
  • GPU, accelerator, HPC, or distributed-compute infrastructure
  • Low-level performance profiling and systems optimization
  • Operating latency-sensitive systems in production.

Strong candidates will be able to explain what they personally optimized, how they measured it, and the production impact it created.

This is an intense, highly hands-on environment with direct founder access and broad ownership. It will suit engineers who want to move across models, kernels, hardware, and infrastructure rather than remain within a narrowly defined area.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Engineering Lead, Inference Optimization
Engineering Lead, Inference Optimization

Shields Group Search • United States

On-site
USD 270,000 - 330,000
Equity
Crypto token compensation
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Inference Engineer
Inference Engineer

Hyperbolic Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000