Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave

San Francisco (CA)

On-site

USD 188,000 - 275,000

Full time

7 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical insurance
Life Insurance
Disability insurance
Tuition Reimbursement
ESPP
Paid Parental Leave
Flexible PTO
Catered lunch

Job summary

CoreWeave is hiring an Applied AI Engineer to understand, measure, and improve the real-world performance of our inference platform. In the near term, you will build and run rigorous benchmarks, profile model and system behavior, identify bottlenecks, and drive targeted optimizations for both platform-wide and customer-specific workloads.

This role focuses on applied performance work within the Inference team, with opportunities to broaden as the product matures and the team expands.

Qualifications

  • 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.
  • Strong programming skills in Python and comfort in production engineering environments.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks.

Responsibilities

  • Build and maintain benchmarking workflows measuring latency, throughput, quality regressions, and cost across priority models and serving configurations.
  • Benchmark the inference stack against realistic customer workloads to identify performance gaps and improvement opportunities.
  • Profile model-serving behavior across frameworks, runtimes, and hardware configurations to find bottlenecks in prefill, decode, KV cache usage, batching, graph capture, quantization, and related systems.
  • Drive targeted optimization efforts for specific customer and product workloads, including tuning serving configurations and validating changes against traces and benchmarks.
  • Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching strategies, routing, and other inference optimizations.

Skills

Python
Performance engineering
Empirical benchmarking
LLM inference systems
Technical communication

Tools

vLLM
SGLang
TensorRT-LLM
Nsight Systems
PyTorch profilers

Job description

CoreWeave is hiring an Applied AI Engineer to understand, measure, and improve the real-world performance of our inference platform. In the near term, you will build and run rigorous benchmarks, profile model and system behavior, identify bottlenecks, and drive targeted optimizations for both platform-wide and customer-specific workloads.

This role focuses on applied performance work within the Inference team, with opportunities to broaden as the product matures and the team expands.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

TheDataJob • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical, dental, vision insurance
ESPP – Employee Stock Purchase Program
Tuition Reimbursement
+3
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Senior AI Infrastructure Performance & Observability Engineer
Senior AI Infrastructure Performance & Observability Engineer

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Health insurance
Life Insurance
Disability insurance
+5
Senior Engineering Manager – AI Inference Platform
Senior Engineering Manager – AI Inference Platform

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 303,000
Medical, dental, vision
Equity awards
ESPP
+2
Senior AI Inference Platform Engineer — Benchmarking
Senior AI Inference Platform Engineer — Benchmarking

Apple • Seattle (WA)

On-site
USD 175,000 - 309,000
Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
+4
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
AI Inference Benchmark Architect
AI Inference Benchmark Architect

Neurophos • Austin (TX), Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Health insurance
Unlimited PTO
401k matching
+3
Distributed Inference Performance Engineer
Distributed Inference Performance Engineer

OpenAI • California (MO)

On-site
USD 150,000 - 190,000
Applied AI Engineer, Inference
Applied AI Engineer, Inference

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
Applied AI Engineer, Inference
Applied AI Engineer, Inference

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1