Engineering Lead, Inference Optimization

Shields Group Search

United States

On-site

USD 270,000 - 330,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Crypto token compensation

Job summary

Shields Group Search is partnering with a privacy‑focused consumer AI company to hire an Engineering Lead for Inference Optimization. This role blends hands‑on work with people leadership to shape the company’s inference stack and scale GPU performance.

You’ll drive latency improvements, build benchmarks, and guide a small team of engineers. Strong GPU, Python (plus Rust/Go) and LLM inference experience are expected, with equity and crypto token compensation alongside a base salary of

Qualifications

  • 8+ years in performance optimization or HPC, with deep GPU architecture knowledge.
  • 5+ years of experience leading engineering teams.
  • Proficiency in Python, Rust, or Go.
  • Hands-on experience with production LLM inference engines (e.g. vLLM, SGLang).
  • Experience with inference optimization techniques and CUDA/Triton tooling.

Responsibilities

  • Own the company’s technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team.
  • Optimize GPU infrastructure across architectures (e.g. H200s, B300s).
  • Improve latency, throughput and cost per token for LLM inference.
  • Build benchmarking harnesses to compare engines and quantization.
  • Work with the routing system to optimize load balancing.
  • Evaluate new techniques and hardware for viability.

Skills

Performance optimization
GPU architecture
Parallel programming
Team leadership
Python
Rust
Go
LLM inference engines
Latency optimization
Profiling tools

Tools

Nsight Systems
Nsight Compute
PyTorch Profiler
Torch.compile
CUDA
Triton kernels
vLLM
SGLang

Job description

Engineering Lead, Inference Optimization
Remote — US Only
Base Salary: $270,000–$330,000 USD + Equity + Crypto Token Compensation

Shields Group Search is working with a leading consumer AI company built on the principles of privacy, free speech, and user sovereignty to hire an Engineering Lead, Inference Optimization.

About the Company

Our client is the world’s leading consumer AI company built on principles of privacy, free speech, and user sovereignty.

They’re building the Port City of AI, in which millions of individuals, third-party apps, and AI agents gather, interact, and access sophisticated AI resources on a private and permissive foundation.

Their mission is to make artificial intelligence approachable and useful in everyday work—bridging the gap between cutting-edge research and practical, real-world impact.

They’re a fast-moving startup where every team member is expected to make a clear impact. Their culture is rooted in curiosity, ownership, ethical principle, philosophy, and collaboration—whether they’re designing better AI workflows, supporting their growing community, or shaping the future of human-AI interaction.

Joining the company means joining a team of unorthodox builders who believe in moving quickly, delivering a beautiful, mass-market, highly useful consumer product that doesn’t spy on people or censor their ideas and questions, and maintaining an edge in the rapidly evolving world of agentic machine intelligence.

If you’re energized by big ideas, entrepreneurial spirit, individual empowerment, and the opportunity to help shape a fast-growing company in the world’s hottest industry from the ground up, you’ll feel right at home here.

Why They’re Hiring

Our client is the only AI platform that runs inference with zero data retention and zero training on user inputs.

This is an opportunity to be on the bleeding edge of privacy-focused AI with a unique and dedicated team of high-agency individuals alongside you.

This role requires both hands-on work as an individual contributor as well as the management of a small team. You will play a pivotal role, shaping the company’s overarching technical strategy and assembling an exceptional team to deliver peak inference performance at massive scale.

The base annual salary for this position ranges from $270,000–$330,000 USD, with equity and crypto token compensation included, and reports to the Head of Engineering.

What You’ll Do
  • Own the company’s technical strategy for inference performance
  • Recruit and lead the Inference Optimization Team
  • Optimize the company’s GPU infrastructure across a range of architectures (e.g. H200s, B300s)
  • Improve latency, throughput, and cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Work with the company’s inference routing system to optimize multivariate inference load-balancing algorithms
  • Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements. Hands-on kernel development experience is a strong plus.
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the company’s stack
Who You Are
  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge
  • 5+ years of experience leading engineering teams
  • Proficiency in Python, Rust, or Go
  • Hands-on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume in production
  • Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
  • Fluency with quantization tradeoffs, both qualitative and quantitative
  • Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments
  • Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler) and a bias toward measuring before optimizing
  • Bonus: diffusion/image model inference optimization, custom Triton kernels, contributions to open-source inference frameworks
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization in San Francisco
Machine Learning Engineer, LLM Inference Optimization in San Francisco

Energy Jobline ZR • San Francisco (CA)

On-site
USD 180,000 - 300,000
Remote Engineering Lead, Inference Optimization
Remote Engineering Lead, Inference Optimization

Shields Group Search • United States

On-site
USD 270,000 - 330,000
Equity
Crypto token compensation
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer (LLM inference)
Machine Learning Engineer (LLM inference)

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000