Remote Engineering Lead, Inference Optimization

Shields Group Search

United States

On-site

USD 270,000 - 330,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Crypto token compensation

Job summary

Shields Group Search is partnering with a privacy‑focused consumer AI company to hire an Engineering Lead for Inference Optimization. This role blends hands‑on work with people leadership to shape the company’s inference stack and scale GPU performance.

You’ll drive latency improvements, build benchmarks, and guide a small team of engineers. Strong GPU, Python (plus Rust/Go) and LLM inference experience are expected, with equity and crypto token compensation alongside a base salary of

Qualifications

  • 8+ years in performance optimization or HPC, with deep GPU architecture knowledge.
  • 5+ years of experience leading engineering teams.
  • Proficiency in Python, Rust, or Go.
  • Hands-on experience with production LLM inference engines (e.g. vLLM, SGLang).
  • Experience with inference optimization techniques and CUDA/Triton tooling.

Responsibilities

  • Own the company’s technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team.
  • Optimize GPU infrastructure across architectures (e.g. H200s, B300s).
  • Improve latency, throughput and cost per token for LLM inference.
  • Build benchmarking harnesses to compare engines and quantization.
  • Work with the routing system to optimize load balancing.
  • Evaluate new techniques and hardware for viability.

Skills

Performance optimization
GPU architecture
Parallel programming
Team leadership
Python
Rust
Go
LLM inference engines
Latency optimization
Profiling tools

Tools

Nsight Systems
Nsight Compute
PyTorch Profiler
Torch.compile
CUDA
Triton kernels
vLLM
SGLang

Job description

Shields Group Search is partnering with a privacy‑focused consumer AI company to hire an Engineering Lead for Inference Optimization. This role blends hands‑on work with people leadership to shape the company’s inference stack and scale GPU performance.

You’ll drive latency improvements, build benchmarks, and guide a small team of engineers. Strong GPU, Python (plus Rust/Go) and LLM inference experience are expected, with equity and crypto token compensation alongside a base salary of

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Lead, Inference Optimization
Engineering Lead, Inference Optimization

Shields Group Search • United States

On-site
USD 270,000 - 330,000
Equity
Crypto token compensation
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Staff Inference Engineer
Staff Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 170,000 - 230,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Remote Inference Optimization Engineer
Remote Inference Optimization Engineer

Modular Mailing Systems, Inc. • Los Altos (CA)

Hybrid
USD 198,000 - 286,000
Premier insurance plans
5% 401k matching
Flexible paid time off
+2