Inference Performance Engineer — Applied AI

CoreWeave

Seattle (WA)

On-site

USD 188,000 - 275,000

Full time

22 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Equity awards
Discretionary bonus
401(k) with match
Flexible PTO
Tuition reimbursement
Parental leave
Catered lunch
Casual work environment

Job summary

CoreWeave is The Essential Cloud for AI. We are looking for an Applied AI Engineer to help us understand, measure, and improve the real-world performance of our inference platform.

In the near term you will build and run rigorous benchmarks, profile model and system behavior, identify bottlenecks, and drive targeted optimizations for platform-wide and customer-specific workloads. You will work with inference platform engineers to productionize improvements, design experiments on quantization,

Qualifications

  • 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.
  • Strong programming skills in Python and comfort working in production engineering environments.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks.

Responsibilities

  • Build and maintain benchmarking workflows that measure latency, throughput, quality regressions, and cost across priority models and serving configurations.
  • Benchmark our inference stack against realistic customer workloads and external provider baselines to identify performance gaps and improvement opportunities.
  • Profile model-serving behavior across frameworks, runtimes, and hardware configurations to find bottlenecks in prefill, decode, KV cache usage, batching, graph capture, quantization, and related systems.
  • Drive targeted optimization efforts for specific customer and product workloads, including tuning serving configurations, evaluating runtime features, and validating changes against representative traces and benchmarks.
  • Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching strategies, routing, and other inference optimizations, with careful attention to quality and correctness tradeoffs.
  • Partner closely with inference platform engineers to productionize improvements and establish repeatable workflows for performance testing and regression detection.
  • Produce clear technical writeups and recommendations that help the team make better decisions about model configurations, runtime choices, hardware allocation, and customer-specific deployment strategies.
  • Contribute additional applied research over time as needed to support inference quality, optimization, and product performance goals.

Skills

Python
Benchmarking
Model serving
Performance tuning
LLM inference

Tools

Nsight Systems
PyTorch profilers
Telemetry pipelines
Benchmark suites

Job description

CoreWeave is The Essential Cloud for AI. We are looking for an Applied AI Engineer to help us understand, measure, and improve the real-world performance of our inference platform.

In the near term you will build and run rigorous benchmarks, profile model and system behavior, identify bottlenecks, and drive targeted optimizations for platform-wide and customer-specific workloads. You will work with inference platform engineers to productionize improvements, design experiments on quantization,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

TheDataJob • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical, dental, vision insurance
ESPP – Employee Stock Purchase Program
Tuition Reimbursement
+3
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Staff AI Inference Systems Engineer
Staff AI Inference Systems Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Voluntary supplemental life insurance
+3
Senior Engineering Manager – AI Inference Platform
Senior Engineering Manager – AI Inference Platform

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 303,000
Medical, dental, vision
Equity awards
ESPP
+2
AI Inference Platform Performance Lead
AI Inference Platform Performance Lead

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 309,000
Medical and dental coverage
Retirement benefits
Discounted products and free services
+1
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Healthcare coverage
Equity awards
401(k) match
+5
Senior Engineering Manager, AI Inference Platform
Senior Engineering Manager, AI Inference Platform

CoreWeave • Bellevue (WA)

On-site
USD 162,000 - 198,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
AI Benchmarking & Performance Architect
AI Benchmarking & Performance Architect

CoreWeave • Sunnyvale (CA)

On-site
USD 206,000 - 333,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+2
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000