Applied AI Inference Engineer - Benchmark & Optimize

CoreWeave

Bellevue (WA)

On-site

USD 188,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
Life Insurance
Disability Insurance
Flexible Spending Account
Health Savings Account
Tuition Reimbursement
ESPP
Mental Wellness Benefits
Parental Leave
Childcare Support
401(k) Matching
Flexible PTO
Catered Lunch

Job summary

CoreWeave is seeking an Applied AI Engineer to advance the real-world performance of our inference platform. You will build benchmarks, profile model-serving behavior, and drive targeted optimizations for customer workloads across frameworks and hardware.

You will run experiments on quantization, decoding, and routing while collaborating with inference platform engineers to productionize improvements. The role emphasizes applied performance work with measurable impact.

Qualifications

  • 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.
  • Strong programming skills in Python and production engineering environments.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks.
  • Understanding of latency, throughput, batching, GPU utilization, quantization, and quality regression analysis.

Responsibilities

  • Build and maintain benchmarking workflows that measure latency, throughput, quality regressions, and cost across priority models and serving configurations.
  • Benchmark our inference stack against realistic customer workloads and external provider baselines to identify performance gaps and improvement opportunities.
  • Profile model-serving behavior across frameworks, runtimes, and hardware configurations to find bottlenecks in prefill, decode, KV cache usage, batching, graph capture, quantization, and related systems.
  • Drive targeted optimization efforts for specific customer and product workloads, including tuning serving configurations, evaluating runtime features, and validating changes against representative traces and benchmarks.
  • Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching strategies, routing, and other inference optimizations, with careful attention to quality and correctness tradeoffs.
  • Partner closely with inference platform engineers to productionize improvements and establish repeatable workflows for performance testing and regression detection.
  • Produce clear technical writeups and recommendations that help the team make better decisions about model configurations, runtime choices, hardware allocation, and customer-specific deployment strategies.
  • Contribute additional applied research over time as needed to support inference quality, optimization, and product performance goals.

Skills

Python
Benchmarks
Profiling
LLM Inference
Experiments

Tools

Nsight Systems
PyTorch Profiler
Telemetry Pipelines
SGLang
TensorRT-LLM

Job description

CoreWeave is seeking an Applied AI Engineer to advance the real-world performance of our inference platform. You will build benchmarks, profile model-serving behavior, and drive targeted optimizations for customer workloads across frameworks and hardware.

You will run experiments on quantization, decoding, and routing while collaborating with inference platform engineers to productionize improvements. The role emphasizes applied performance work with measurable impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Inference Performance Engineer
Inference Performance Engineer

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
+2
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical benefits
401(k) with match
Flexible PTO
+5
Senior Performance Engineer for AI Platform
Senior Performance Engineer for AI Platform

Weights & Biases • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
AI Performance & Benchmarking Tech Lead
AI Performance & Benchmarking Tech Lead

Coreweave • United States

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+5
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3