LLM Inference Library Engineer - High-Performance AI

Jobot

San Francisco (CA)

On-site

USD 175,000 - 250,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity (startup)
Competitive compensation
Healthcare, vision, dental

Job summary

Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from complex hardware.

We value experience in AI/ML infrastructure, CUDA/Rocm, and model-serving frameworks, with strong Python and C++/Rust skills. Equity, competitive pay, and excellent benefits are offered.

Qualifications

  • Experience in AI/ML infrastructure or HPC software.
  • Programming in Python and C++, Rust, or similar systems languages.
  • Experience with LLM inference frameworks and model-serving infrastructure.
  • Hands-on CUDA, ROCm, Triton or similar GPU programming.
  • Understanding of transformer/LLM architectures and memory management.
  • Experience benchmarking AI workloads across hardware environments.
  • Familiarity with batching, attention, KV caching, quantization, and memory management.

Responsibilities

  • Build and maintain a high-performance inference library for AI models.
  • Optimize performance across diverse compute architectures.
  • Work with CUDA/ROCm and related frameworks to accelerate workloads.
  • Analyze and improve batching, KV caching, and memory usage.

Skills

AI/ML infrastructure
Python
C++
Rust
LLM inference
GPU programming
Performance benchmarking
Transformer/LLM architectures
Memory management
Kernel optimization

Tools

CUDA
ROCm
Triton
vLLM
TensorRT-LLM
SGLang

Job description

Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from complex hardware.

We value experience in AI/ML infrastructure, CUDA/Rocm, and model-serving frameworks, with strong Python and C++/Rust skills. Equity, competitive pay, and excellent benefits are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Global Inference Library Engineer
Global Inference Library Engineer

Jobot • San Francisco (CA)

On-site
USD 175,000 - 250,000
Equity (startup)
Competitive compensation
Healthcare, vision, dental
Global Inference Library Engineer
Global Inference Library Engineer

LeoForce • San Francisco (CA)

On-site
USD 175,000 - 250,000
Healthcare
Vision care
Dental
+1
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference Library Engineer - GPU-Optimized
Senior AI Inference Library Engineer - GPU-Optimized

LeoForce • San Francisco (CA)

On-site
USD 175,000 - 250,000
Healthcare
Vision care
Dental
+1
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Distributed LLM Inference & Optimization Engineer
Distributed LLM Inference & Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup equity
Health insurance
Competitive benefits
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
LLM Inference Architect & Systems Optimizer
LLM Inference Architect & Systems Optimizer

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits