GPU Kernel Engineer — High-Performance ML at Scale

The Consensus

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Flexible PTO policy including a Winter Break
Paid parental leave
Fertility and family-building stipend

Job summary

The Consensus is looking for a GPU Kernel Engineer to optimize machine learning performance. The ideal candidate will design high-performance GPU kernels and collaborate on cutting-edge projects in the AI field. This role offers substantial growth opportunities in an inclusive environment.

You'll be a critical part of shaping AI applications, working with state-of-the-art tools, and contributing to significant advancements in GPU performance.

Qualifications

  • Proficient in optimizing code using CUDA and PTX assembly.
  • Knowledge of memory access patterns and bandwidth optimization.
  • Experience with GPU kernel libraries such as Cutlass and Triton.

Responsibilities

  • Design and implement high-performance GPU kernels for ML operations.
  • Collaborate with research teams to productionize advancements.
  • Contribute to internal and open-source GPU libraries.

Skills

Strong understanding of GPU architecture
Proficient in C++
Experience with GPU performance profiling tools

Tools

CUDA C++ API
Nsight Systems
Torch Profiler

Job description

The Consensus is looking for a GPU Kernel Engineer to optimize machine learning performance. The ideal candidate will design high-performance GPU kernels and collaborate on cutting-edge projects in the AI field. This role offers substantial growth opportunities in an inclusive environment.

You'll be a critical part of shaping AI applications, working with state-of-the-art tools, and contributing to significant advancements in GPU performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Senior GPU Kernel & Performance Engineer
Senior GPU Kernel & Performance Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
Dental insurance
Vision insurance
+2
CUDA Kernel Engineer — Optimize GPU Performance at Scale
CUDA Kernel Engineer — Optimize GPU Performance at Scale

Pragmatike • California (MO)

On-site
USD 180,000 - 240,000
Health, Dental, and Vision
Sign-on bonus
401k
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
GPU Kernel Optimization Engineer for AI Training
GPU Kernel Optimization Engineer for AI Training

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
Fellow GPU Kernel Engineer - ML/HPC Optimizations
Fellow GPU Kernel Engineer - ML/HPC Optimizations

AMD • Austin (TX)

Hybrid
USD 190,000 - 270,000
GPU Kernel Performance Engineer (Contract)
GPU Kernel Performance Engineer (Contract)

Obsidian • San Francisco (CA)

Remote
USD 83,000 - 138,000
GPU Systems Research Intern: High-Performance ML Kernels
GPU Systems Research Intern: High-Performance ML Kernels

Together Computer Inc • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Software Engineer - GPU Kernel
Software Engineer - GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3