Applied AI Inference Performance Engineer

CoreWeave

San Francisco (CA)

On-site

USD 188,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
401(k) with match
Flexible PTO
Catered lunch
Tuition Reimbursement
Employee Stock Purchase Program (ESPP)
Parental Leave
Family-forming support

Job summary

CoreWeave, a leading AI cloud provider, seeks an Applied AI Engineer to improve real-world performance of the inference platform. You will benchmark, profile, and optimize across frameworks, runtimes, and GPUs, delivering measurable improvements for customer workloads.

The role focuses on applied performance work, including quantization, decoding, and caching strategies, with close collaboration with the Inference Team. Strong Python skills and 4+ years’ experience are required.

Qualifications

  • 4+ years of experience in ML, systems, performance engineering, or related applied engineering work.
  • Strong Python programming and production engineering experience.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar stacks.
  • Understanding practical tradeoffs in latency, throughput, batching, GPU utilization, quantization, and quality regression analysis.
  • Ability to work across model, systems, and product boundaries with customer outcomes in mind.
  • Strong written communication and reproducibility skills.

Responsibilities

  • Build and maintain benchmarking workflows measuring latency, throughput, quality regressions, and cost across models and serving configurations.
  • Benchmark inference stack against realistic customer workloads and external baselines to identify gaps.
  • Profile model-serving behavior across frameworks, runtimes, and hardware to find bottlenecks.
  • Drive targeted optimization for specific workloads, tuning configurations and validating changes with traces and benchmarks.
  • Design experiments on quantization, speculative decoding, caching, routing, and other inference optimizations with quality tradeoffs.
  • Collaborate with inference platform engineers to productionize improvements and establish repeatable performance testing workflows.
  • Produce technical writeups and recommendations for model configurations, runtimes, hardware allocations, and deployment strategies.
  • Contribute applied research to support inference quality, optimization, and product performance goals.

Skills

ML/Systems
Python
Benchmarking
LLM inference
Performance analysis
Cross-functional collaboration

Tools

Nsight Systems
PyTorch profilers

Job description

CoreWeave, a leading AI cloud provider, seeks an Applied AI Engineer to improve real-world performance of the inference platform. You will benchmark, profile, and optimize across frameworks, runtimes, and GPUs, delivering measurable improvements for customer workloads.

The role focuses on applied performance work, including quantization, decoding, and caching strategies, with close collaboration with the Inference Team. Strong Python skills and 4+ years’ experience are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Inference Engineer - Benchmark & Optimize
Applied AI Inference Engineer - Benchmark & Optimize

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical Insurance
Dental Insurance
Vision Insurance
+12
Inference Performance Engineer
Inference Performance Engineer

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
+2
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
AI Inference Engineer - Performance & API
AI Inference Engineer - Performance & API

Topazlabs • Dallas (TX)

On-site
USD 120,000 - 180,000
Full medical/dental/vision coverage
15 days PTO
5 personal days + holidays
+1
Cloud GPU Performance Engineer for AI Infra
Cloud GPU Performance Engineer for AI Infra

Coreweave • United States

On-site
USD 120,000 - 160,000
Health insurance
Life insurance
Disability insurance
+11
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
Senior Performance Engineer for AI Platform
Senior Performance Engineer for AI Platform

Weights & Biases • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3