Senior GPU Kernel Engineer: Inference Throughput

CoreWeave

Bellevue (WA)

On-site

USD 182,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
Tuition Reimbursement
Flexible PTO
Catered lunch in office

Job summary

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel authoring and optimization for LLM inference. You will write and tune CUDA kernels, optimize tensor cores, and push end-to-end latency down while maintaining accuracy.

You will lead benchmarking workflows (MLPerf), mentor engineers, and collaborate with cross-functional partners. Experience with CUDA, C++, Python, and GPU architectures is required; familiarity with vLLM, TensorRT-LLM, llm-d, SGLang is a

Qualifications

  • 5+ years of experience in HPC, GPU software, or performance-critical systems.
  • Hands-on CUDA experience with written/optimized kernels and CUDA memory model.
  • Deep understanding of GPU architecture, tensor cores, warp occupancy, memory hierarchy.

Responsibilities

  • Author, profile, and optimize CUDA kernels for LLM inference on the critical path.
  • Tune tensor cores, occupancy, memory coalescing, shared memory, and data movement overlap.
  • Benchmark end-to-end latency/throughput; maintain reproducible MLPerf workflows and runbooks.

Skills

High-performance computing
CUDA experience
C++
Python
GPU architecture
Model-serving stacks
Cross-functional collaboration

Tools

Nsight Compute
Nsight Systems
NCCL
Kubernetes

Job description

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel authoring and optimization for LLM inference. You will write and tune CUDA kernels, optimize tensor cores, and push end-to-end latency down while maintaining accuracy.

You will lead benchmarking workflows (MLPerf), mentor engineers, and collaborate with cross-functional partners. Experience with CUDA, C++, Python, and GPU architectures is required; familiarity with vLLM, TensorRT-LLM, llm-d, SGLang is a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer: Inference Performance
Senior GPU Kernel Engineer: Inference Performance

Neura Market • Sunnyvale (CA), Northern (KY)

Hybrid
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+2
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior GPU Kernel Optimizer for LLM Inference
Senior GPU Kernel Optimizer for LLM Inference

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000