Senior AI Inference Library Engineer - GPU-Optimized

LeoForce

San Francisco (CA)

On-site

USD 175,000 - 250,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare
Vision care
Dental
Equity in seed-stage startup

Job summary

LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures, squeezing maximum performance while integrating with model-serving infrastructure.

The role emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building cutting-edge AI software.

Qualifications

  • Strong experience building AI/ML infrastructure, inference systems, or HPC software.
  • Strong programming experience with Python and C++, Rust, or similar systems languages.
  • Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies.
  • Experience developing, integrating, or optimizing performance-critical compute kernels.

Responsibilities

  • Build and maintain a high-performance inference library for modern AI models.
  • Support inferencing across diverse compute architectures and accelerators.
  • Profile, optimize, and tune performance-critical kernels and workflows.
  • Collaborate on model-serving infrastructure and related tooling for scalability.

Skills

AI infrastructure
Python
C++
Rust
GPU programming
LLM inference
Performance profiling

Tools

vLLM
TensorRT-LLM
SGLang
CUDA
Triton

Job description

LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures, squeezing maximum performance while integrating with model-serving infrastructure.

The role emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building cutting-edge AI software.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Global Inference Library Engineer
Global Inference Library Engineer

LeoForce • San Francisco (CA)

On-site
USD 175,000 - 250,000
Healthcare
Vision care
Dental
+1
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
LLM Inference Library Engineer - High-Performance AI
LLM Inference Library Engineer - High-Performance AI

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
Global Inference Library Engineer
Global Inference Library Engineer

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
GPU Inference Engineer — Deep Learning
GPU Inference Engineer — Deep Learning

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits