Global Inference Library Engineer

LeoForce

San Francisco (CA)

On-site

USD 175,000 - 250,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare
Vision care
Dental
Equity in seed-stage startup

Job summary

LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures, squeezing maximum performance while integrating with model-serving infrastructure.

The role emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building cutting-edge AI software.

Qualifications

  • Strong experience building AI/ML infrastructure, inference systems, or HPC software.
  • Strong programming experience with Python and C++, Rust, or similar systems languages.
  • Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies.
  • Experience developing, integrating, or optimizing performance-critical compute kernels.

Responsibilities

  • Build and maintain a high-performance inference library for modern AI models.
  • Support inferencing across diverse compute architectures and accelerators.
  • Profile, optimize, and tune performance-critical kernels and workflows.
  • Collaborate on model-serving infrastructure and related tooling for scalability.

Skills

AI infrastructure
Python
C++
Rust
GPU programming
LLM inference
Performance profiling

Tools

vLLM
TensorRT-LLM
SGLang
CUDA
Triton

Job description

Global Inference Library Engineer

Experience: Senior Level

Salary: $175,000 - $250,000 per year

Job Details

We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments.

What We’re Looking For
  • Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software
  • Strong programming experience with Python and C++, Rust, or similar systems languages
  • Experience with LLM inference frameworks and model-serving infrastructure
  • Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies
  • Experience developing, integrating, or optimizing performance-critical compute kernels
  • Understanding of modern transformer and LLM architectures
  • Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management
  • Experience benchmarking and profiling AI workloads across different hardware environments
  • Strong understanding of GPU or accelerator architecture and performance characteristics
  • Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable
A bit about us:

We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures.

Why join us?
  • Well-funded by leading tech investors
  • Cutting edge technical problems with complex solutions
  • Lucrative Equity in a seed stage startup
  • Competitive compensation
  • Excellent benefits (healthcare, vision, dental)

#techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Global Inference Library Engineer
Global Inference Library Engineer

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior AI Inference Library Engineer - GPU-Optimized
Senior AI Inference Library Engineer - GPU-Optimized

LeoForce • San Francisco (CA)

On-site
USD 175,000 - 250,000
Healthcare
Vision care
Dental
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
LLM Inference Library Engineer - High-Performance AI
LLM Inference Library Engineer - High-Performance AI

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer – AI Inference Engine
Software Engineer – AI Inference Engine

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance