Engineering Lead, Inference Optimization New Remote- US only

Venice.ai, Inc.

Northern (KY)

Hybrid

USD 270,000 - 330,000

Full time

10 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Venice AI, Inc. is seeking an Engineering Lead to drive inference performance and scalable GPU optimization for private, high-volume AI workloads.

You will build and lead the Inference Optimization Team, optimize across architectures, and push latency, throughput, and cost per token to new levels while shaping the technical strategy for Venice’s private AI stack. Reporting to the Head of Engineering, the role combines hands-on development with people leadership in a fast-moving startup

Qualifications

  • 8+ years in performance optimization or HPC.
  • 5+ years leading engineering teams.
  • Proficiency in Python, Rust, or Go.
  • Hands-on experience with a production LLM inference engine (e.g. vLLM, SGLang).
  • Demonstrated experience with LLM inference optimization techniques: batching, KV cache management, quantization, CUDA graphs.

Responsibilities

  • Own Venice’s technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team at Venice.
  • Optimize Venice's GPU infrastructure across architectures (e.g. H200s, B300s).
  • Improve latency, throughput, and cost per token for LLM inference workloads.
  • Build reproducible benchmarking harnesses across inference engines to identify the optimal engine and strategy per workload.
  • Work with our inference routing system to optimize multivariate inference load-balancing algorithms.
  • Evaluate emerging inference optimization techniques and hardware viability for Venice's stack.

Skills

GPU optimization
Python
Rust
Go
HPC
Leadership
Production LLM
CUDA

Tools

vLLM
SGLang
CUDA
Triton
Nsight
torch.compile

Job description

Engineering Lead, Inference Optimization

Remote- US only

About Us

Venice is the world’s leading consumer AI company built on principles of privacy, free speech, and user sovereignty.

We’re building the Port City of AI, in which millions of individuals, third party apps, and AI agents gather, interact, and access sophisticated AI resources on a private and permissive foundation. Our mission is to make artificial intelligence approachable and useful in everyday work—bridging the gap between cutting-edge research and practical, real-world impact.

We’re a fast-moving startup where every team member is expected to make a clear impact. Our culture is rooted in curiosity, ownership, ethical principle, philosophy, and collaboration—whether we’re designing better AI workflows, supporting our growing community, or shaping the future of human-AI interaction.

Joining Venice AI means joining a team of unorthodox builders who believe in moving quickly, delivering a beautiful mass-marketand highly useful consumer product that doesn’t spy on people or censor their ideas and questions., and maintaining an edge in the rapidly evolving world of agentic machine intelligence. If you’re energized by big ideas, entrepreneurial spirit, individual empowerment, and the opportunity to help shape a fast‑growing company in the world’s hottest industry from the ground up.

Why we are hiring

Venice is the only AI platform that runs inference with zero data retention and zero training on user inputs. This is an opportunity for you to be on the bleeding edge of privacy-focused AI with a unique and dedicated team of high‑agency individuals alongside you. This role requires both hands‑on work as an individual contributor as well as the management of a small team. You will play a pivotal role, shaping Venice's overarching technical strategy and assembling an exceptional team to deliver peak inference performance at massive scale.

The base annual salary for this position ranges from $270,000-$330,000 USD and reports to the Head of Engineering.

What you’ll do
  • Own Venice’s technical strategy for inference performance
  • Recruit and lead the Inference Optimization Team at Venice
  • Optimize Venice's GPU infrastructure across a range of architectures (e.g. H200s, B300s)
  • Improve latency, throughput, and cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Work with our inference routing system to optimize multivariate inference load‑balancing algorithms
  • Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements. Hands‑on kernel development experience is a strong plus.
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in Venice's stack.
Who you are
  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge
  • 5+ years experience leading engineering teams
  • Proficiency in Python, Rust, or Go.
  • Hands‑on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume in production
  • Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
  • Fluency with quantization tradeoffs, both qualitative and quantitative
  • Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi‑GPU and multi‑node environments
  • Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler) and a bias toward measuring before optimizing
  • Bonus: diffusion/image model inference optimization, custom Triton kernels, contributions to open‑source inference frameworks
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Engineering Lead, Inference Optimization
Remote Engineering Lead, Inference Optimization

Venice.ai, Inc. • Northern (KY)

Hybrid
USD 270,000 - 330,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

Remote
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Santa Clara (CA)

Hybrid
USD 224,000 - 431,250
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits package
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package