Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave

Sunnyvale (CA)

On-site

USD 182,000 - 242,000

Full time

14 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Equity awards
401(k) match
Flexible PTO
Tuition reimbursement
ESPP
Parental leave
Lunch provided

Job summary

CoreWeave, the Essential Cloud for AI, is seeking a Senior Engineer for the Benchmarking & Performance team in Sunnyvale, CA. You will write, profile, and tune CUDA kernels on the critical path of large-scale model serving to maximize throughput and minimize latency.

You will lead end-to-end benchmarking efforts including MLPerf submissions, collaborate with product, orchestration, and hardware teams, and mentor junior engineers while raising coding and testing standards.

Qualifications

  • 5+ years of experience in HPC, GPU software, or performance-critical systems.
  • Hands-on CUDA experience with custom kernel development and optimization.
  • Deep understanding of GPU architecture, tensor cores, memory hierarchy, and profiling.
  • Strong coding in C++ and Python; reading/writing low-level performance-sensitive code.
  • Familiarity with model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang).

Responsibilities

  • Author, profile, and optimize CUDA kernels on the LLM inference path.
  • Tune occupancy, memory coalescing, and data movement for max throughput and min latency.
  • Develop benchmarking workflows and reproduce results for MLPerf inferences and training.
  • Lead design reviews, decompose multi-service work into milestones, and raise standards.
  • Mentor junior engineers and ensure reproducible benchmarking practices.

Skills

CUDA
C++
Python
GPU-architecture
Nsight Compute

Education

Bachelor's degree in CS/Math/Engineering

Tools

CUDA toolkit
Nsight Compute
TensorRT

Job description

CoreWeave, the Essential Cloud for AI, is seeking a Senior Engineer for the Benchmarking & Performance team in Sunnyvale, CA. You will write, profile, and tune CUDA kernels on the critical path of large-scale model serving to maximize throughput and minimize latency.

You will lead end-to-end benchmarking efforts including MLPerf submissions, collaborate with product, orchestration, and hardware teams, and mentor junior engineers while raising coding and testing standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior CUDA Kernel Engineer for High-Performance Inference
Senior CUDA Kernel Engineer for High-Performance Inference

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+3
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
AI Benchmarking & Performance Architect
AI Benchmarking & Performance Architect

CoreWeave • Sunnyvale (CA)

On-site
USD 206,000 - 333,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+2
Senior GPU Systems Performance Architect
Senior GPU Systems Performance Architect

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Senior GPU Benchmarking & Optimization Engineer
Senior GPU Benchmarking & Optimization Engineer

Webhosting • United States

On-site
USD 140,000 - 150,000
Health insurance
401(k) plan with matching
Professional development reimbursement
+4
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000