Inference Performance Engineer

Neura Market

Bellevue, Northern (WA, KY)

Hybrid

USD 188,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
Flexible PTO
Catered lunch

Job summary

CoreWeave is seeking an Applied AI Engineer to understand and boost the real-world performance of the inference platform. You will build benchmarks, profile model-serving behavior, and drive targeted optimizations for production workloads.

The role emphasizes empirical evaluation, reproducible experiments, and close collaboration with inference platform engineers to productionize improvements and guide hardware and configuration choices.

Qualifications

  • 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.
  • Strong programming skills in Python and comfort working in production engineering environments.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks.

Responsibilities

  • Build and maintain benchmarking workflows that measure latency, throughput, quality regressions, and cost across priority models and serving configurations.
  • Benchmark our inference stack against realistic customer workloads and external provider baselines to identify performance gaps and improvement opportunities.
  • Profile model-serving behavior across frameworks, runtimes, and hardware configurations to find bottlenecks in prefill, decode, KV cache usage, batching, graph capture, quantization, and related systems.
  • Drive targeted optimization efforts for specific customer and product workloads, including tuning serving configurations, evaluating runtime features, and validating changes against representative traces and benchmarks.
  • Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching strategies, routing, and other inference optimizations, with careful attention to quality and correctness tradeoffs.
  • Partner closely with inference platform engineers to productionize improvements and establish repeatable workflows for performance testing and regression detection.
  • Produce clear technical writeups and recommendations that help the team make better decisions about model configurations, runtime choices, hardware allocation, and customer-specific deployment strategies.
  • Contribute additional applied research over time as needed to support inference quality, optimization, and product performance goals.

Skills

Python
ML engineering
Benchmarks
Profiling
LLM inference

Tools

Nsight Systems
PyTorch Profiler
Telemetry pipelines

Job description

CoreWeave is seeking an Applied AI Engineer to understand and boost the real-world performance of the inference platform. You will build benchmarks, profile model-serving behavior, and drive targeted optimizations for production workloads.

The role emphasizes empirical evaluation, reproducible experiments, and close collaboration with inference platform engineers to productionize improvements and guide hardware and configuration choices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Applied AI Inference Engineer - Benchmark & Optimize
Applied AI Inference Engineer - Benchmark & Optimize

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical Insurance
Dental Insurance
Vision Insurance
+12
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical benefits
401(k) with match
Flexible PTO
+5
Senior Performance Engineer for AI Platform
Senior Performance Engineer for AI Platform

Weights & Biases • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
AI Inference Platform Performance Engineer
AI Inference Platform Performance Engineer

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior AI/ML Platform Engineering Manager (Inference)
Senior AI/ML Platform Engineering Manager (Inference)

CoreWeave • New York (NY)

On-site
USD 188,000 - 303,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
+3