Staff AI Inference Engineer — Production Optimizations

Crusoe

United States

On-site

USD 215,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Restricted Stock Units
Paid time off, paid holidays & leave
Health insurance (comprehensive)
HSA contributions
Paid parental leave
Life insurance and disability
Professional development & tuition
Mental health & wellness support
Commuter benefits
Cell phone stipend
401(k) Retirement plan with company 4%
Volunteer time off
Travel insurance & emergency support
Daily meals allowance
Location-specific perks

Job summary

Crusoe is seeking a hands-on engineer to make large language models run faster, cheaper, and more reliably in production. You will own the inference stack end to end, from profiling time and cost to optimizing CUDA kernels and deployment strategies.

You will work with customer teams to tailor deployments, ship production-ready improvements, and drive performance across diverse models using Python as a primary language.

Qualifications

  • Bachelor's, Master's, or Ph.D. in CS/Engineering/Math or related field.
  • Hands-on production code experience in Python or C++ (prefer Python).
  • Experience optimizing LLMs for high throughput / low latency.
  • Familiar with LLM serving frameworks such as vLLM or SGLang; profiling to kernel level.
  • Strong understanding of GPU architecture and behavior.
  • Interest and hands-on experience with large language models.
  • Knowledge of AI/ML pipelines from development to deployment.
  • Strong communication skills for explaining technical topics.

Responsibilities

  • Bring inference techniques into production and refine them.
  • Design and optimize serving architectures for latency, throughput, and cost.
  • Tune the serving stack from frameworks to kernels, profiling for performance.
  • Adapt optimization methods across many ML models, focusing on LLMs.
  • Profile deployments against latency targets and ensure reliability.
  • Collaborate with customer engineering teams to tailor deployments.
  • Build and support inference stack software in production using Python/C++.
  • Experiment quickly and ship well-tested results.
  • Deliver end-to-end optimizations from initial experiments to production.
  • Work through ambiguity and make sound trade-off calls.
  • Take ownership and accountability in the role.

Skills

Python
C++
LLM inference optimization
Performance profiling
GPU basics

Education

Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field

Tools

vLLM
SGLang
CUDA
Docker
Kubernetes

Job description

Crusoe is seeking a hands-on engineer to make large language models run faster, cheaper, and more reliably in production. You will own the inference stack end to end, from profiling time and cost to optimizing CUDA kernels and deployment strategies.

You will work with customer teams to tailor deployments, ship production-ready improvements, and drive performance across diverse models using Python as a primary language.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production AI Inference Engineer for Fast LLMs
Production AI Inference Engineer for Fast LLMs

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+6
Senior AI Inference Engineer - Production LLM Optimizer
Senior AI Inference Engineer - Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe • Denver (CO)

On-site
USD 185,000 - 225,000
Equity packages
RSUs
Health insurance
+12
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe • United States

On-site
USD 215,000 - 260,000
Competitive compensation
Restricted Stock Units
Paid time off, paid holidays & leave
+13
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+6
Senior AI Cloud Platform Engineer - Reliability & Scale
Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Senior AI Infra Engineer - Scalable LLM Platforms
Senior AI Infra Engineer - Scalable LLM Platforms

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 205,000
Health insurance
401(k) with company match
Paid parental leave
+6
Senior AI Inference Customer Success Manager
Senior AI Inference Customer Success Manager

Crusoe • Denver (CO)

On-site
USD 190,000 - 215,000
Competitive compensation and equity
Restricted Stock Units
Paid time off
+6
Senior AI Customer Success Leader for Production Inference
Senior AI Customer Success Leader for Production Inference

Crusoe • San Francisco (CA)

On-site
USD 190,000 - 215,000
Competitive compensation
Equity packages
Paid time off and holidays
+6