GPU Kernel Engineer for High-Performance AI Inference

Baseten

New York (NY)

On-site

USD 180,000 - 360,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
Fertility and family-building stipend
Company-facilitated 401(k)

Job summary

Baseten is seeking a GPU Kernel Engineer in New York, NY. The ideal candidate will design high-performance GPU kernels and optimize computation for AI workloads. Responsibilities include implementing advanced features, contributing to open-source GPU libraries, and collaborating with research teams. The position offers competitive compensation, 100% coverage of insurance, and flexible PTO. Join Baseten to be part of cutting-edge AI acceleration efforts.

Qualifications

  • Deep understanding of GPU memory hierarchy and thread organization.
  • Experience optimizing GPU code for performance.
  • Knowledge of memory access patterns and thread synchronization.

Responsibilities

  • Design high-performance GPU kernels for ML operations.
  • Optimize code using CUDA and architecture-specific techniques.
  • Resolve performance bottlenecks in GPU applications.

Skills

Strong understanding of GPU architecture
Proficient in C++
Knowledge of CUDA C++ API
Experience with numerical precision
Familiarity with performance profiling tools

Tools

CUDA
Nsight Systems
Torch Profiler
Cutlass
Triton

Job description

Baseten is seeking a GPU Kernel Engineer in New York, NY. The ideal candidate will design high-performance GPU kernels and optimize computation for AI workloads. Responsibilities include implementing advanced features, contributing to open-source GPU libraries, and collaborating with research teams. The position offers competitive compensation, 100% coverage of insurance, and flexible PTO. Join Baseten to be part of cutting-edge AI acceleration efforts.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
+2
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Senior GPU Inference Engine Engineer
Senior GPU Inference Engine Engineer

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Flexible PTO policy including a Winter Break
+2
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,900 - 232,800
Annual performance bonus
Equity options
Comprehensive benefits package