Senior GPU Kernel Engineer: Inference Performance

Neura Market

Sunnyvale, Northern (CA, KY)

Hybrid

USD 182,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
Employee Stock Purchase Program (ESPP)
Flexible PTO

Job summary

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements.

You will lead kernel design reviews, mentor engineers, and drive reproducible benchmarking across vLLM, TensorRT-LLM, and related stacks while partnering with product, orchestration, and hardware teams.

Qualifications

  • 5+ years of experience building high-performance computing, GPU/accelerator software, or performance-critical systems.
  • Hands-on CUDA experience is required—you have written and optimized custom kernels and are fluent with the CUDA programming and memory model.
  • Deep understanding of GPU architecture and performance: tensor cores, warp/occupancy tuning, the memory hierarchy and bandwidth, NVLink/PCIe, and profiling with Nsight Compute/Systems.
  • Strong coding in C++ and Python; comfortable reading and writing low-level, performance-sensitive code.
  • Familiarity with model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang) and the kernels that dominate their inference cost.
  • Strong communicator comfortable collaborating with cross-functional teams and external partners.

Responsibilities

  • Author, profile, and optimize CUDA kernels—GEMMs, attention, MoE routing, quantization, KV-cache, and fused epilogues—on the critical path of LLM inference.
  • Optimize for the hardware: exploit tensor cores and tune occupancy, memory coalescing, shared-memory/register usage, and overlap of compute with data movement.
  • Use kernel-authoring DSLs and compilers to prototype and ship kernels quickly without sacrificing performance.
  • Benchmark rigorously: build reproducible microbenchmarks and roofline analyses, and validate that kernel-level wins translate to end-to-end latency/throughput gains across model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang).
  • Implement and maintain benchmarking workflows for end-to-end MLPerf Inference (and Training) runs, including workload setup, cluster configuration, runbooks, and result validation.
  • Lead design reviews and drive architecture within the team; decompose multi-service work into clear milestones.
  • Mentor junior engineers; review cross-team designs and elevate coding/testing standards.
  • Help ensure reproducible, well-documented benchmarking and kernel-optimization processes.

Skills

CUDA kernels
C++
Python
GPU architecture
Performance optimization
Mentoring
Cross-functional collaboration

Tools

Nsight Compute
TensorRT

Job description

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements.

You will lead kernel design reviews, mentor engineers, and drive reproducible benchmarking across vLLM, TensorRT-LLM, and related stacks while partnering with product, orchestration, and hardware teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer: Inference Throughput
Senior GPU Kernel Engineer: Inference Throughput

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+3
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior GPU Kernel Optimizer for LLM Inference
Senior GPU Kernel Optimizer for LLM Inference

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000