Inference Optimization Lead

Up Top

United States

Hybrid

USD 180,000 - 320,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Up Top seeks a hybrid IC/leader to own the technical strategy for large-scale inference performance and to build and guide a dedicated Inference Optimization team.

You will drive latency reduction, throughput improvements, and cost-per-token optimization across a modern GPU fleet, coordinating with benchmarking and routing across multiple inference engines and emerging hardware platforms.

Qualifications

  • 8+ years in performance optimization or HPC
  • 5+ years leading engineering teams
  • Proficiency in Python, Rust, or Go
  • Hands-on experience running production LLM inference engines at high volume
  • Depth in modern inference optimization: continuous batching, KV-cache management, speculative decoding, quantization, CUDA graphs, and torch.compile

Responsibilities

  • Own the end-to-end technical strategy for inference performance across the platform
  • Recruit, build, and lead the Inference Optimization team
  • Optimize GPU infrastructure across current and next-gen architectures (e.g. H200, B300)
  • Improve latency, throughput, and cost-per-token for production LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines such as vLLM and SGLang
  • Tune inference routing and multivariate load-balancing algorithms
  • Evaluate emerging optimization techniques — custom CUDA/Triton kernels, attention variants, quantization schemes, and compilation improvements
  • Assess emerging inference hardware (FPGAs, ASICs, and custom silicon)

Skills

Performance optimization
Team leadership
Python / Rust / Go
GPU profiling
LLM inference experience

Tools

Nsight Systems
Nsight Compute
PyTorch Profiler
CUDA
torch.compile

Job description

Up Top seeks a hybrid IC/leader to own the technical strategy for large-scale inference performance and to build and guide a dedicated Inference Optimization team.

You will drive latency reduction, throughput improvements, and cost-per-token optimization across a modern GPU fleet, coordinating with benchmarking and routing across multiple inference engines and emerging hardware platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Lead Inference Performance Architect
Lead Inference Performance Architect

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health benefits
Unlimited PTO
401(k) matching
+3
Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Hands-On Tech Lead for Large-Scale Inference
Hands-On Tech Lead for Large-Scale Inference

Luma AI • United States

Remote
USD 180,000 - 320,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior Tech Lead, On-Device AI Inference
Senior Tech Lead, On-Device AI Inference

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Lead, AI Compute Infra for Scalable LLM Inference
Lead, AI Compute Infra for Scalable LLM Inference

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500