Inference Optimization Engineer

US Health Partners, LLC

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
Lunch and snacks at the office

Job summary

Hedra is a small, highly technical team in San Francisco building state-of-the-art visual models and the infrastructure to deploy them at scale. We seek an Inference Optimization Engineer to push model performance at inference time, working at the boundary between research and systems.

You’ll collaborate with researchers to run new architectures efficiently on modern hardware, exploring novel optimization techniques and scalable execution across accelerators.

Qualifications

  • Deep technical ability in efficient ML inference, ML systems, GPU computing.
  • Strong programming fundamentals in Python, C++, or another systems language.
  • Experience with PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies.
  • Ability to reason across abstraction layers and collaborate across research and engineering boundaries.

Responsibilities

  • Work with research scientists and engineers to optimize inference for new visual models.
  • Profile architectures and workloads to identify bottlenecks in compute, memory, and communication.
  • Develop and implement approaches to reduce latency and improve throughput.
  • Explore optimizations including quantization, sparsity, caching, and alternative execution strategies.
  • Build or optimize GPU kernels using CUDA, Triton, or similar tools.
  • Optimize execution across single-GPU, multi-GPU, and multi-node setups.
  • Study interaction between model design and hardware to unlock performance gains.
  • Create benchmarking and profiling infrastructure to evaluate progress.
  • Evaluate new runtimes, compilers, and accelerator hardware.

Skills

Python
C++
ML systems
GPU computing
Profiling ML workloads

Tools

PyTorch
CUDA
Triton
TensorRT
vLLM

Job description

About Hedra

Hedra is the platform, models, and infrastructure for visual intelligence.

We build models and systems that push the frontier of visual intelligence, along with the infrastructure required to make those models fast, efficient, reliable, and accessible at scale.

We’re a small, highly technical team in San Francisco, backed by a16z and other leading investors. Researchers and engineers at Hedra work closely across boundaries, own problems end to end, and have significant influence over both what we build and how we build it.

The Role

We’re looking for an Inference Optimization Engineer to work alongside our research team on making state-of-the-art visual models fast and efficient at inference time.

You’ll work at the boundary between research and systems, taking new model architectures and figuring out how to run them efficiently on modern hardware. That means understanding where time and memory are being spent, identifying opportunities for algorithmic and systems-level improvements, and implementing optimizations across model architecture, inference algorithms, runtimes, kernels, and distributed execution.

The problems rarely live neatly within one layer of the stack. Depending on what you find, you might modify how a model executes, develop a new inference technique, write a custom GPU kernel, rethink memory movement, or change how work is distributed across accelerators.

We care more about technical depth, curiosity, and demonstrated ability than years of experience. We’re open to experienced ML systems engineers as well as exceptional early-career engineers or researchers who have already gone unusually deep on model performance, GPU systems, or efficient inference.

What You’ll Do
  • Work directly with research scientists and engineers to make new visual models fast and efficient at inference time.

  • Profile model architectures and workloads to understand bottlenecks across compute, memory, communication, and model execution.

  • Develop and implement new approaches to improving inference latency, throughput, memory efficiency, and GPU utilization.

  • Explore algorithmic optimizations including quantization, sparsity, caching, compilation, attention optimizations, and alternative execution strategies.

  • Build or optimize GPU kernels using CUDA, Triton, or similar technologies when existing implementations leave performance on the table.

  • Optimize model execution across single-GPU, multi-GPU, and multi-node environments.

  • Reason about the interaction between model architecture and hardware, and work with researchers when architectural changes can unlock meaningful performance improvements.

  • Investigate communication, memory movement, parallelism, and distributed execution strategies for large visual models.

  • Build rigorous benchmarking, profiling, and performance-regression infrastructure to understand performance and evaluate new optimization ideas.

  • Evaluate new inference runtimes, compilers, frameworks, optimization techniques, and accelerator hardware.

  • Stay close to advances in efficient inference, GPU programming, model architectures, compilers, and ML systems research, and rapidly test promising ideas.

  • Help turn research breakthroughs into models that can be deployed and served efficiently at scale.

What We’re Looking For
  • Deep technical ability in efficient ML inference, ML systems, GPU computing, or adjacent research, demonstrated through research, production engineering, open-source contributions, or unusually ambitious independent work.

  • Strong understanding of how modern deep learning models execute on hardware, including the relationship between compute, memory, communication, and performance.

  • Experience profiling ML workloads, identifying bottlenecks, forming hypotheses, and driving measurable performance improvements.

  • Strong programming fundamentals in Python, C++, or another systems-oriented language.

  • Experience with some combination of PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies.

  • Ability to reason across abstraction layers rather than treating model architecture, framework, runtime, kernel, and hardware boundaries as fixed.

  • Strong intuition for performance tradeoffs across latency, throughput, memory, numerical precision, model quality, and complexity.

  • Curiosity about how models work internally and a willingness to modify or rethink existing approaches when the performance problem calls for it.

  • Comfort working on ambiguous problems where the bottleneck, and sometimes even the right question, is not known in advance.

  • Ability to communicate technical ideas clearly and collaborate closely with research scientists and engineers.

We don’t expect every candidate to have experience across the entire stack. Exceptional depth in one or more relevant areas, combined with the ability and curiosity to reason across the others, matters more to us than checking every box.

Nice to Have
  • Experience optimizing large generative, multimodal, vision, or video models.

  • CUDA, Triton, CUTLASS, or other GPU kernel development.

  • Deep knowledge of GPU architecture, memory hierarchy, and hardware-aware optimization.

  • Experience with attention optimization, kernel fusion, memory-efficient execution, or custom operators.

  • Model compilation or graph optimization experience.

  • Quantization, sparsity, caching, speculative execution, or other efficient inference techniques.

  • Experience optimizing diffusion, autoregressive, transformer, or other large generative architectures.

  • Multi-GPU or multi-node model execution, including tensor, pipeline, sequence, or other forms of parallelism.

  • Experience optimizing communication or data movement between accelerators.

  • Experience with profiling tools such as Nsight Systems or Nsight Compute.

  • Contributions to ML systems, inference runtimes, compilers, GPU libraries, or performance-focused open-source projects.

  • Research or publications in efficient ML, ML systems, GPU computing, compilers, or related areas.

Benefits:
  • Competitive compensation and equity

  • 401k

  • Healthcare (Silver PPO Medical, Vision, Dental)

  • Lunch and snacks at the office

This role is based in San Francisco, and we work together in person five days a week.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Optimization Engineer
Inference Optimization Engineer

Speedrun Talent Network • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Inference Optimization Engineer
Inference Optimization Engineer

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
401k
Healthcare
+1
Senior Modeling Architect, Performance Benchmarking
Senior Modeling Architect, Performance Benchmarking

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

On-site
USD 150,000 - 210,000
Health benefits
Unlimited PTO
401(k) matching
+3
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

On-site
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior/Staff Software Engineer, Distributed Systems
Senior/Staff Software Engineer, Distributed Systems

Hedra Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Competitive compensation and equity
401k
Healthcare (Silver PPO Medical, Vision
+1
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Modeling Architect
Modeling Architect

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

On-site
USD 150,000 - 230,000
Health coverage
HSA contributions
Unlimited PTO
+3
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily