Remote Engineering Lead, Inference Optimization

Venice.ai, Inc.

Northern (KY)

Hybrid

USD 270,000 - 330,000

Full time

30 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Venice AI, Inc. is seeking an Engineering Lead to drive inference performance and scalable GPU optimization for private, high-volume AI workloads.

You will build and lead the Inference Optimization Team, optimize across architectures, and push latency, throughput, and cost per token to new levels while shaping the technical strategy for Venice’s private AI stack. Reporting to the Head of Engineering, the role combines hands-on development with people leadership in a fast-moving startup

Qualifications

  • 8+ years in performance optimization or HPC.
  • 5+ years leading engineering teams.
  • Proficiency in Python, Rust, or Go.
  • Hands-on experience with a production LLM inference engine (e.g. vLLM, SGLang).
  • Demonstrated experience with LLM inference optimization techniques: batching, KV cache management, quantization, CUDA graphs.

Responsibilities

  • Own Venice’s technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team at Venice.
  • Optimize Venice's GPU infrastructure across architectures (e.g. H200s, B300s).
  • Improve latency, throughput, and cost per token for LLM inference workloads.
  • Build reproducible benchmarking harnesses across inference engines to identify the optimal engine and strategy per workload.
  • Work with our inference routing system to optimize multivariate inference load-balancing algorithms.
  • Evaluate emerging inference optimization techniques and hardware viability for Venice's stack.

Skills

GPU optimization
Python
Rust
Go
HPC
Leadership
Production LLM
CUDA

Tools

vLLM
SGLang
CUDA
Triton
Nsight
torch.compile

Job description

Venice AI, Inc. is seeking an Engineering Lead to drive inference performance and scalable GPU optimization for private, high-volume AI workloads.

You will build and lead the Inference Optimization Team, optimize across architectures, and push latency, throughput, and cost per token to new levels while shaping the technical strategy for Venice’s private AI stack. Reporting to the Head of Engineering, the role combines hands-on development with people leadership in a fast-moving startup

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Lead, Inference Optimization New Remote- US only
Engineering Lead, Inference Optimization New Remote- US only

Venice.ai, Inc. • Northern (KY)

Hybrid
USD 270,000 - 330,000
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Senior Remote LLM Inference Optimization Lead
Senior Remote LLM Inference Optimization Lead

Dragonfly Digital Management, LLC (Dragonfly Capital) • United States

On-site
USD 140,000 - 210,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Distributed AI Inference Performance Engineer
Distributed AI Inference Performance Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model