INFERENCE OPTIMIZATION ENGINEER

Up Top

United States

Hybrid

USD 180,000 - 320,000

Full time

25 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Up Top seeks a hybrid IC/leader to own the technical strategy for large-scale inference performance and to build and guide a dedicated Inference Optimization team.

You will drive latency reduction, throughput improvements, and cost-per-token optimization across a modern GPU fleet, coordinating with benchmarking and routing across multiple inference engines and emerging hardware platforms.

Qualifications

  • 8+ years in performance optimization or HPC
  • 5+ years leading engineering teams
  • Proficiency in Python, Rust, or Go
  • Hands-on experience running production LLM inference engines at high volume
  • Depth in modern inference optimization: continuous batching, KV-cache management, speculative decoding, quantization, CUDA graphs, and torch.compile

Responsibilities

  • Own the end-to-end technical strategy for inference performance across the platform
  • Recruit, build, and lead the Inference Optimization team
  • Optimize GPU infrastructure across current and next-gen architectures (e.g. H200, B300)
  • Improve latency, throughput, and cost-per-token for production LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines such as vLLM and SGLang
  • Tune inference routing and multivariate load-balancing algorithms
  • Evaluate emerging optimization techniques — custom CUDA/Triton kernels, attention variants, quantization schemes, and compilation improvements
  • Assess emerging inference hardware (FPGAs, ASICs, and custom silicon)

Skills

Performance optimization
Team leadership
Python / Rust / Go
GPU profiling
LLM inference experience

Tools

Nsight Systems
Nsight Compute
PyTorch Profiler
CUDA
torch.compile

Job description

Our client is a fast-growing consumer AI company built around privacy and user ownership. Their platform lets a large and growing base of individuals, third-party apps, and AI agents interact on a privacy-preserving foundation — with no retention of user data and no training on user inputs. The team is small, fast-moving, and product-obsessed, with a culture rooted in curiosity, ownership, and collaboration.

ABOUT THE ROLE

This is a hybrid IC / leadership role reporting to the Head of Engineering. You'll own the technical strategy for inference performance at scale, then build and lead a dedicated Inference Optimization team. The mandate is simple to state and hard to execute: drive down latency, increase throughput, and cut cost-per-token for large-scale LLM inference across a modern GPU fleet.

WHAT YOU'LL DO
  • Own the end-to-end technical strategy for inference performance across the platform
  • Recruit, build, and lead the Inference Optimization team
  • Optimize GPU infrastructure across current and next-gen architectures (e.g. H200, B300)
  • Improve latency, throughput, and cost-per-token for production LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines such as vLLM and SGLang
  • Tune inference routing and multivariate load-balancing algorithms
  • Evaluate emerging optimization techniques — custom CUDA/Triton kernels, attention variants, quantization schemes, and compilation improvements
  • Assess emerging inference hardware (FPGAs, ASICs, and custom silicon)
WHAT WE'RE LOOKING FOR
  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel-programming expertise
  • 5+ years leading engineering teams
  • Proficiency in Python, Rust, or Go
  • Hands-on experience running production LLM inference engines at high volume
  • Depth in modern inference optimization: continuous batching, PagedAttention / KV-cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
  • A strong grasp of quantization tradeoffs
  • Experience with distributed inference across multi-GPU / multi-node environments
  • Fluency with GPU profiling tools (Nsight Systems, Nsight Compute, PyTorch Profiler)
BONUS POINTS
  • C++ / CUDA proficiency
  • Diffusion / image-model inference optimization
  • Contributions to open-source inference frameworks
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Lead, Inference Optimization
Engineering Lead, Inference Optimization

Shields Group Search • United States

On-site
USD 270,000 - 330,000
Equity
Crypto token compensation
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer (LLM inference)
Machine Learning Engineer (LLM inference)

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000