Fast, Scalable Vision Inference Engineer

US Health Partners, LLC

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
Lunch and snacks at the office

Job summary

Hedra is a small, highly technical team in San Francisco building state-of-the-art visual models and the infrastructure to deploy them at scale. We seek an Inference Optimization Engineer to push model performance at inference time, working at the boundary between research and systems.

You’ll collaborate with researchers to run new architectures efficiently on modern hardware, exploring novel optimization techniques and scalable execution across accelerators.

Qualifications

  • Deep technical ability in efficient ML inference, ML systems, GPU computing.
  • Strong programming fundamentals in Python, C++, or another systems language.
  • Experience with PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies.
  • Ability to reason across abstraction layers and collaborate across research and engineering boundaries.

Responsibilities

  • Work with research scientists and engineers to optimize inference for new visual models.
  • Profile architectures and workloads to identify bottlenecks in compute, memory, and communication.
  • Develop and implement approaches to reduce latency and improve throughput.
  • Explore optimizations including quantization, sparsity, caching, and alternative execution strategies.
  • Build or optimize GPU kernels using CUDA, Triton, or similar tools.
  • Optimize execution across single-GPU, multi-GPU, and multi-node setups.
  • Study interaction between model design and hardware to unlock performance gains.
  • Create benchmarking and profiling infrastructure to evaluate progress.
  • Evaluate new runtimes, compilers, and accelerator hardware.

Skills

Python
C++
ML systems
GPU computing
Profiling ML workloads

Tools

PyTorch
CUDA
Triton
TensorRT
vLLM

Job description

Hedra is a small, highly technical team in San Francisco building state-of-the-art visual models and the infrastructure to deploy them at scale. We seek an Inference Optimization Engineer to push model performance at inference time, working at the boundary between research and systems.

You’ll collaborate with researchers to run new architectures efficiently on modern hardware, exploring novel optimization techniques and scalable execution across accelerators.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Performance Engineer for Visual AI
Inference Performance Engineer for Visual AI

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
401k
Healthcare
+1
Inference Optimization Engineer
Inference Optimization Engineer

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
401k
Healthcare
+1
Inference Optimization Engineer
Inference Optimization Engineer

Speedrun Talent Network • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Inference Optimization Engineer
Inference Optimization Engineer

US Health Partners, LLC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Senior/Staff Software Engineer, Distributed Systems
Senior/Staff Software Engineer, Distributed Systems

Hedra Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Senior Distributed Systems Engineer for Inference Platform
Senior Distributed Systems Engineer for Inference Platform

Hedra Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
AI/ML Inference & Vision Infrastructure Engineer
AI/ML Inference & Vision Infrastructure Engineer

Frontdoor Defense • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Restaurant d'entreprise
Indemnités de stage/alternance
Performance Engineer, Inference Engine - Flexible Hours
Performance Engineer, Inference Engine - Flexible Hours

Anthropic • San Francisco (CA), New York (NY)

On-site
USD 350,000 - 850,000
Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

adaption • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work in Bay Area
Adaption Passport travel stipend
Lunch stipend
+1