Inference Performance Engineer for Visual AI

AI Chopping Block, Inc.

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
401k
Healthcare
Lunch & snacks

Job summary

Hedra in San Francisco is seeking an Inference Optimization Engineer to join our research-and-systems team. You will push state-of-the-art visual models to run faster at inference time across single and multi-GPU environments.

Role spans algorithmic and systems optimizations, including quantization and advanced memory techniques, with close collaboration with researchers and engineers to deploy scalable solutions.

Qualifications

  • Deep technical ability in efficient ML inference, ML systems, GPU computing, or adjacent research.
  • Strong understanding of how modern deep learning models execute on hardware.
  • Experience profiling ML workloads, identifying bottlenecks, and driving improvements.
  • Strong programming fundamentals in Python, C++, or another systems language.
  • Experience with PyTorch, CUDA, Triton, TensorRT, vLLM, or comparable technologies.

Responsibilities

  • Work with researchers and engineers to make new visual models fast and efficient at inference time.
  • Profile architectures and workloads to identify bottlenecks in compute, memory, and communication.
  • Develop approaches to reduce latency, increase throughput, and improve memory efficiency.
  • Explore algorithmic optimizations including quantization, sparsity, and memory movement.
  • Build or optimize GPU kernels using CUDA or Triton when needed.
  • Optimize execution across single- and multi-GPU, and multi-node setups.
  • Reason about architecture-hardware interactions to unlock performance gains.
  • Develop benchmarking and profiling infrastructure to evaluate ideas.
  • Evaluate new runtimes, compilers, and accelerator hardware.

Skills

Python
C++
GPU computing
Profiling ML workloads
PyTorch
CUDA
Triton
TensorRT
vLLM
SGLang

Tools

Nsight Systems
Nsight Compute
CUTLASS
GPU kernel development

Job description

Hedra in San Francisco is seeking an Inference Optimization Engineer to join our research-and-systems team. You will push state-of-the-art visual models to run faster at inference time across single and multi-GPU environments.

Role spans algorithmic and systems optimizations, including quantization and advanced memory techniques, with close collaboration with researchers and engineers to deploy scalable solutions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Fast, Scalable Vision Inference Engineer
Fast, Scalable Vision Inference Engineer

US Health Partners, LLC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

adaption • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work in Bay Area
Adaption Passport travel stipend
Lunch stipend
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI, Inc. • San Francisco (CA)

On-site
USD 295,000 - 555,000
Equity
Head of Inference Performance & System Visibility
Head of Inference Performance & System Visibility

Etched • San Jose (CA)

On-site
USD 240,000 - 340,000
Medical, dental, and vision
Housing subsidy
Relocation support
+3
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
AI/ML Inference & Vision Infrastructure Engineer
AI/ML Inference & Vision Infrastructure Engineer

Frontdoor Defense • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Restaurant d'entreprise
Indemnités de stage/alternance
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Real-Time GPU Optimization Engineer - Inference
Real-Time GPU Optimization Engineer - Inference

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000