Senior AI Inference Engineer: High-Throughput LLMs

Crusoe Energy Systems

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits
401(k) match
Paid time off
Mental wellness support

Job summary

Crusoe Energy Systems in San Francisco is hiring an ML systems engineer focused on making large language models run faster, cheaper, and more reliably in production. You will own the end-to-end inference stack, profiling time and cost, and bringing modern optimization techniques to real deployments.

The role involves hands-on coding, low-level optimization, and collaboration with customer engineering teams to tailor deployments to their workloads and latency targets.

Qualifications

  • Proven experience shipping production-grade ML inference systems.
  • Strong background in optimizing LLMs for latency and throughput.
  • Experience profiling kernels and serving stacks for performance.
  • Comfort with customer-facing discussions and requirements.

Responsibilities

  • Own end-to-end optimization of inference stacks for evolving models.
  • Profile, tune, and debug performance across GPUs and kernels.
  • Collaborate with customers to tailor deployments and monitor results.
  • Lead experiments from concept to production with measurable gains.

Skills

Production code
LLM optimization
Performance profiling
Customer-facing
Python
Communication skills
AI/ML pipelines

Education

Bachelor/Master/PhD in CS or related

Tools

CUDA
Docker
Kubernetes
vLLM
SGLang

Job description

Crusoe Energy Systems in San Francisco is hiring an ML systems engineer focused on making large language models run faster, cheaper, and more reliably in production. You will own the end-to-end inference stack, profiling time and cost, and bringing modern optimization techniques to real deployments.

The role involves hands-on coding, low-level optimization, and collaboration with customer engineering teams to tailor deployments to their workloads and latency targets.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Engineer - Production LLM Optimizer
Senior AI Inference Engineer - Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Production AI Inference Engineer for Fast LLMs
Production AI Inference Engineer for Fast LLMs

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+6
Senior AI Infra Engineer - Scalable LLM Platforms
Senior AI Infra Engineer - Scalable LLM Platforms

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 205,000
Health insurance
401(k) with company match
Paid parental leave
+6
Staff AI Inference Engineer — Production Optimizations
Staff AI Inference Engineer — Production Optimizations

Crusoe • United States

On-site
USD 215,000 - 260,000
Competitive compensation
Restricted Stock Units
Paid time off, paid holidays & leave
+13
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Senior AI Model Lifecycle Engineer - LLM Training Pipelines
Senior AI Model Lifecycle Engineer - LLM Training Pipelines

AI • San Francisco (CA)

On-site
USD 237,000 - 319,000
Competitive compensation
Restricted Stock Units
Paid time off & holidays
+4
Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Engineering Manager, AI Platform & LLM Infra
Engineering Manager, AI Platform & LLM Infra

Crusoe • United States

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+7