Inference Performance Engineer - Benchmark & Optimize

CoreWeave

Sunnyvale (CA)

On-site

USD 188,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
Flexible PTO
Equity options

Job summary

CoreWeave, The Essential Cloud for AI, is seeking an Applied AI Engineer to improve the real-world performance of our inference platform. You will build benchmarks, profile model-serving pipelines, identify bottlenecks, and drive targeted optimizations across frameworks, runtimes, and hardware.

This role emphasizes empirical evaluation, quantization, caching, and deployment strategies, with opportunities to collaborate with inference engineers to productionize changes and deliver measurable

Qualifications

  • 4+ years of ML, systems, performance engineering or adjacent applied engineering work.
  • Strong programming skills in Python and production engineering familiarity.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM.

Responsibilities

  • Build and maintain benchmarking workflows measuring latency, throughput, quality regressions, and cost across models and serving configurations.
  • Benchmark the inference stack against realistic workloads to find gaps and improvements.
  • Profile model-serving behavior across frameworks, runtimes, and hardware to identify bottlenecks in prefill, decode, KV cache, batching, and quantization.
  • Drive targeted optimization efforts for customer workloads, including tuning serving configs and evaluating runtime features.

Skills

Python
ML performance
Benchmarks
Clear communication

Tools

vLLM
SGLang
TensorRT-LLM
Nsight Systems
PyTorch profilers

Job description

CoreWeave, The Essential Cloud for AI, is seeking an Applied AI Engineer to improve the real-world performance of our inference platform. You will build benchmarks, profile model-serving pipelines, identify bottlenecks, and drive targeted optimizations across frameworks, runtimes, and hardware.

This role emphasizes empirical evaluation, quantization, caching, and deployment strategies, with opportunities to collaborate with inference engineers to productionize changes and deliver measurable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Inference Performance Engineer
Inference Performance Engineer

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
+2
Applied AI Inference Engineer - Benchmark & Optimize
Applied AI Inference Engineer - Benchmark & Optimize

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical Insurance
Dental Insurance
Vision Insurance
+12
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical benefits
401(k) with match
Flexible PTO
+5
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior Performance Engineer for AI Platform
Senior Performance Engineer for AI Platform

Weights & Biases • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
AI Performance & Benchmarking Tech Lead
AI Performance & Benchmarking Tech Lead

Coreweave • United States

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+5
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy