Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave

United States

On-site

USD 182,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance
401(k) Matching
Paid Parental Leave
Flexible PTO
Equity Awards
Tuition Reimbursement

Job summary

CoreWeave is the top-rated AI-cloud for high-performance GPU infrastructure. We’re seeking a Senior Engineer for Benchmarking & Performance to own kernel design and optimization on large-scale model serving, squeezing throughput and reducing latency across the inference stack.

You will write, profile, and tune CUDA kernels (GEMMs, attention, KV-cache) and lead end-to-end benchmarking such as MLPerf, partnering with product, orchestration, and hardware teams to meet strict P99 SLAs at scale.

Qualifications

  • 5+ years of experience building high-performance computing, GPU/accelerator software, or performance-critical systems.
  • Hands-on CUDA experience with custom kernel development and memory model proficiency.
  • Strong coding in C++ and Python; reading/writing low-level, performance-sensitive code.
  • Familiarity with model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang) and the kernels that dominate their inference cost.
  • Strong communicator capable of collaborating with cross-functional teams and external partners.

Responsibilities

  • Author, profile, and optimize CUDA kernels on the critical path of LLM inference.
  • Tune for hardware performance: tensor cores, occupancy, memory hierarchy, data movement overlap.
  • Benchmark rigor: build reproducible microbenchmarks and end-to-end latency/throughput validation.
  • Lead design reviews and drive architecture; decompose work into milestones.
  • Mentor junior engineers and elevate coding/testing standards.

Skills

CUDA kernels
GPU architecture
C++
Python
Nsight Compute
Model-serving stacks
Kubernetes
MLPerf
OSS contributions
KNYFE / DSLs

Tools

Nsight Compute/Systems
TensorRT
Kubernetes (prod)

Job description

CoreWeave is the top-rated AI-cloud for high-performance GPU infrastructure. We’re seeking a Senior Engineer for Benchmarking & Performance to own kernel design and optimization on large-scale model serving, squeezing throughput and reducing latency across the inference stack.

You will write, profile, and tune CUDA kernels (GEMMs, attention, KV-cache) and lead end-to-end benchmarking such as MLPerf, partnering with product, orchestration, and hardware teams to meet strict P99 SLAs at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3
Senior GPU Kernel Engineer: Inference Performance
Senior GPU Kernel Engineer: Inference Performance

Neura Market • Sunnyvale (CA), Northern (KY)

Hybrid
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+2
Senior GPU Kernel Engineer: Inference Throughput
Senior GPU Kernel Engineer: Inference Throughput

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+3
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
AI Performance & Benchmarking Tech Lead
AI Performance & Benchmarking Tech Lead

Coreweave • United States

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+5
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1