Senior AI Inference Engineer — Production-Scale LLMs

AI Chopping Block, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 250,000 - 300,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation and equity
RSUs
Paid time off
Health, dental & vision insurance
401(k) with company match
Parental leave
Professional development
Wellness support
Meal allowance

Job summary

Crusoe is hiring an AI inference engineer to optimize large language model deployments in production. You will own the inference stack end to end, profiling time and cost, and collaborating with customer teams to tailor deployments for latency and throughput targets.

The role emphasizes hands-on coding, low-level optimization, and direct impact on real customer workloads, using Python as a primary language and strong GPU awareness.

Qualifications

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • Hands-on experience shipping code in production with Python or C++.
  • Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
  • Comfort with LLM serving frameworks such as vLLM or SGLang and profiling down to the kernel level.
  • Understanding GPUs and how they behave.

Responsibilities

  • Bring current inference techniques into production and refine them.
  • Design and optimize serving architectures, including prefill, decode disaggregation and routing.
  • Profile and tune deployments against latency, throughput, and cost targets.
  • Adapt and scale optimization methods across many ML models, especially large language models.
  • Own delivery end to end from experiment to production.

Skills

Python
C++
LLM knowledge
Strong communication

Education

Bachelor's/Master's/PhD in CS or related field

Tools

Docker
Kubernetes
CUDA

Job description

Crusoe is hiring an AI inference engineer to optimize large language model deployments in production. You will own the inference stack end to end, profiling time and cost, and collaborating with customer teams to tailor deployments for latency and throughput targets.

The role emphasizes hands-on coding, low-level optimization, and direct impact on real customer workloads, using Python as a primary language and strong GPU awareness.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production AI Inference Engineer for Fast LLMs
Production AI Inference Engineer for Fast LLMs

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+6
Senior AI Inference Architect for Production LLMs
Senior AI Inference Architect for Production LLMs

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+13
Senior AI Inference Engineer - Production LLM Optimizer
Senior AI Inference Engineer - Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Senior AI Inference Engineer — Production LLM Optimizer
Senior AI Inference Engineer — Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation
Equity packages
Health/dental/vision insurance
+1
Staff AI Inference Engineer — Production Optimizations
Staff AI Inference Engineer — Production Optimizations

Crusoe • United States

On-site
USD 215,000 - 260,000
Competitive compensation
Restricted Stock Units
Paid time off, paid holidays & leave
+13
Senior AI Inference Engineer: High-Throughput LLMs
Senior AI Inference Engineer: High-Throughput LLMs

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Senior Engineering Manager, AI Inference (LLM Production)
Senior Engineering Manager, AI Inference (LLM Production)

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+13
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

On-site
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe • Denver (CO)

On-site
USD 185,000 - 225,000
Equity packages
RSUs
Health insurance
+12
Engineering Manager, AI Platform & LLM Infra
Engineering Manager, AI Platform & LLM Infra

Crusoe • United States

On-site
USD 215,000 - 260,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+7