Senior AI Inference Architect for Production LLMs

Crusoe

San Francisco (CA)

On-site

USD 250,000 - 300,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
Health, dental & vision insurance
HSA contributions
Parental leave
Life Insurance
Tuition reimbursement
Mental health & wellness support
Commuter benefits
Cell phone stipend
401(k) match
Volunteer time off
Global travel insurance
Daily meals allowance
Location-specific perks

Job summary

Crusoe is seeking an engineer to accelerate large language models in production. You will own the end-to-end inference stack—profiling, optimizing, and deploying fast, cost-effective models for customer workloads.

Expect hands-on work across Python and C++, CUDA, and frameworks like vLLM/SGLang. You’ll collaborate with customer teams, push performance improvements into production, and shape product requirements with engineering and solutions teams.

Qualifications

  • Shipping production-grade code in Python or C++
  • Experience optimizing LLMs for high throughput/low latency
  • Familiarity with LLM serving frameworks like vLLM or SGLang
  • Understanding GPU architectures and behavior
  • Interest and hands-on work with large language models
  • Knowledge of ML pipelines and deploying ML models
  • Strong ability to explain technical topics to customers and teammates

Responsibilities

  • Bring current inference techniques into production and refine them.
  • Design and optimize serving architectures including prefill, decode disaggregation, and routing.
  • Work deeply in the serving stack from frameworks to CUDA kernels, profiling to fix perf issues.
  • Scale optimization methods across many ML models, focusing on large language models.
  • Profile and tune deployments against latency, throughput, and cost targets.
  • Tailor deployments to each customer’s models and constraints with their engineers.
  • Build and support inference stack features in production using Python and other languages.
  • Run rapid experiments, draft specs, and ship well-tested results.
  • Own delivery end-to-end from experiments to production performance.

Skills

Python
C++
LLM optimization
Profiling
Performance tuning
CUDA
Docker
Kubernetes
Communication

Education

Bachelor’s, Master’s, or Ph.D. in CS/Engineering/Math or related field

Tools

vLLM
SGLang
CUDA kernels

Job description

Crusoe is seeking an engineer to accelerate large language models in production. You will own the end-to-end inference stack—profiling, optimizing, and deploying fast, cost-effective models for customer workloads.

Expect hands-on work across Python and C++, CUDA, and frameworks like vLLM/SGLang. You’ll collaborate with customer teams, push performance improvements into production, and shape product requirements with engineering and solutions teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Engineer — Production-Scale LLMs
Senior AI Inference Engineer — Production-Scale LLMs

AI Chopping Block, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation and equity
RSUs
Paid time off
+6
Production AI Inference Engineer for Fast LLMs
Production AI Inference Engineer for Fast LLMs

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+6
Senior AI Inference Engineer - Production LLM Optimizer
Senior AI Inference Engineer - Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Senior AI Inference Engineer — Production LLM Optimizer
Senior AI Inference Engineer — Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation
Equity packages
Health/dental/vision insurance
+1
Senior AI Inference Engineer: High-Throughput LLMs
Senior AI Inference Engineer: High-Throughput LLMs

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Staff AI Inference Engineer — Production Optimizations
Staff AI Inference Engineer — Production Optimizations

Crusoe • United States

On-site
USD 215,000 - 260,000
Competitive compensation
Restricted Stock Units
Paid time off, paid holidays & leave
+13
Senior Engineering Manager — AI Inference in Production
Senior Engineering Manager — AI Inference in Production

crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation
Equity
Paid time off
+2
Senior Engineering Manager, AI Inference (LLM Production)
Senior Engineering Manager, AI Inference (LLM Production)

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+13
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000