Inference Performance Engineer: Benchmark & Optimize

Coreweave

Bellevue (CA)

On-site

USD 188,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Equity awards
401(k) with match
Paid parental leave

Job summary

CoreWeave is seeking an Applied AI Engineer to advance the performance of its inference platform. You will build benchmarks, profile model-serving behavior, and drive targeted optimizations for production workloads across frameworks, runtimes, and hardware.

You will collaborate with the Inference team to design experiments on quantization, caching, and routing, translating results into concrete engineering decisions and scalable improvements for customers.

Qualifications

  • 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.
  • Strong programming skills in Python and comfort in production engineering environments.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks.
  • Understanding of latency, throughput, batching, GPU utilization, quantization, and quality regression analysis.
  • Ability to work across model, systems, and product boundaries with customer-focused outcomes.
  • Strong written communication and reproducible technical work.

Responsibilities

  • Build benchmarking workflows measuring latency, throughput, quality regressions, and cost across priority models and serving configurations.
  • Benchmark our inference stack against realistic customer workloads and external baselines to identify performance gaps and opportunities.
  • Profile model-serving behavior across frameworks, runtimes, and hardware configurations to find bottlenecks in prefill, decode, KV cache usage, batching, and related systems.
  • Drive targeted optimization efforts for specific customer and product workloads, including tuning serving configurations and evaluating runtime features.
  • Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching, and routing, with quality and correctness tradeoffs.
  • Partner closely with inference platform engineers to productionize improvements and establish repeatable workflows for performance testing and regression detection.
  • Produce clear technical writeups and recommendations that help the team decide on model configurations and deployment strategies.
  • Contribute additional applied research over time to support inference quality, optimization, and product performance goals.

Skills

Python
ML systems
Benchmarks/experiments
LLM inference tools
Latency optimization
Cross-functional collaboration
Technical writing

Tools

vLLM
SGLang
TensorRT-LLM

Job description

CoreWeave is seeking an Applied AI Engineer to advance the performance of its inference platform. You will build benchmarks, profile model-serving behavior, and drive targeted optimizations for production workloads across frameworks, runtimes, and hardware.

You will collaborate with the Inference team to design experiments on quantization, caching, and routing, translating results into concrete engineering decisions and scalable improvements for customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Applied AI Inference Engineer - Benchmark & Optimize
Applied AI Inference Engineer - Benchmark & Optimize

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical Insurance
Dental Insurance
Vision Insurance
+12
Inference Performance Engineer
Inference Performance Engineer

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
+2
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical benefits
401(k) with match
Flexible PTO
+5
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior Performance Engineer for AI Platform
Senior Performance Engineer for AI Platform

Weights & Biases • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
AI Performance & Benchmarking Tech Lead
AI Performance & Benchmarking Tech Lead

Coreweave • United States

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+5
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3