Senior CUDA Kernel Engineer for High-Performance Inference

CoreWeave

Bellevue (WA)

On-site

USD 182,000 - 242,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
401(k) with employer match
Flexible PTO
Catered lunch

Job summary

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to write, profile, and tune CUDA kernels on the inference path. You will enhance latency, throughput, and reliability across the stack and work across product, orchestration, and hardware teams.

You will mentor engineers, lead design reviews, and help deliver industry-leading MLPerf results. A strong CUDA background and GPU architecture expertise are required for this critical role at CoreWeave.

Qualifications

  • 5+ years of experience building high-performance computing, GPU software, or performance-critical systems.
  • Hands-on CUDA experience with custom kernel writing and optimization.
  • Deep understanding of GPU architecture including tensor cores and memory hierarchy.
  • Proficient in C++ and Python; able to read/write low-level, performance-sensitive code.
  • Familiarity with model-serving stacks and kernels that dominate inference cost.

Responsibilities

  • Author, profile, and optimize CUDA kernels on the path of LLM inference.
  • Tune for tensor cores, occupancy, memory coalescing, and data movement overlap.
  • Prototype and ship kernels using DSLs and compilers while maintaining performance.
  • Benchmark end-to-end latency and throughput, validate with reproducible results.
  • Lead design reviews, drive architecture, and mentor junior engineers.
  • Ensure reproducible benchmarking and well-documented processes.

Skills

CUDA programming
C++
Python
GPU architecture
communication

Tools

Nsight Compute/Systems
CUDA kernel development

Job description

CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to write, profile, and tune CUDA kernels on the inference path. You will enhance latency, throughput, and reliability across the stack and work across product, orchestration, and hardware teams.

You will mentor engineers, lead design reviews, and help deliver industry-leading MLPerf results. A strong CUDA background and GPU architecture expertise are required for this critical role at CoreWeave.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Healthcare coverage
Equity awards
401(k) match
+5
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Senior CUDA Kernel & Performance Engineer
Senior CUDA Kernel & Performance Engineer

General Motors • Washington

Hybrid
USD 170,000 - 258,000
Health and wellbeing benefits
Hybrid work option
Competitive compensation package
Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Performance Analytics Lead – Kernel & Platform
Senior Performance Analytics Lead – Kernel & Platform

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 272,000 - 489,000
Equity
Benefits
AI Benchmarking & Performance Architect
AI Benchmarking & Performance Architect

CoreWeave • Sunnyvale (CA)

On-site
USD 206,000 - 333,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+2