Kernel Engineer: GPU Performance & Inference

Acceler8 Talent

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco seeks a Member of Technical Staff focusing on Kernels & GPU Performance to push the limits of production AI inference across diverse accelerators.

You will implement low-level kernels, analyze memory hierarchies, and collaborate with compiler, ML systems, and runtime teams to optimize latency and throughput.

This on-site role offers a full-time path at a fast-growing AI infrastructure company.

Qualifications

  • Bachelor’s degree in a relevant technical discipline or equivalent practical experience.
  • Strong fundamentals in software engineering and performance.
  • Experience with low-level execution, memory hierarchies, and scheduling.
  • Ability to profile, debug, and optimize latency and throughput.
  • Experience across multiple accelerator architectures.
  • Familiarity with CUDA/Triton/CUTLASS is a plus.

Responsibilities

  • Build and optimize kernels for production AI workloads.
  • Improve latency, throughput, and hardware utilization.
  • Develop execution strategies across multiple accelerator architectures.
  • Optimize memory efficiency, scheduling behavior, and low-level execution characteristics.
  • Analyze accelerator execution models and memory hierarchies to identify bottlenecks.
  • Profile and validate performance across different hardware platforms.
  • Partner with compiler, runtime, ML systems, and distributed-systems engineers on end-to-end performance optimization.
  • Develop optimization approaches that account for architectural differences between accelerators.
  • Help establish performance-engineering standards and best practices across the execution platform.

Skills

Software-engineering fundamentals
Performance-critical systems
Low-level execution
Memory hierarchy understanding
Profiling & debugging

Education

Bachelor’s degree in a relevant technical discipline

Tools

CUDA
Triton
CUTLASS
Accelerator programming models
GPU profiling tools

Job description

Acceler8 Talent in San Francisco seeks a Member of Technical Staff focusing on Kernels & GPU Performance to push the limits of production AI inference across diverse accelerators.

You will implement low-level kernels, analyze memory hierarchies, and collaborate with compiler, ML systems, and runtime teams to optimize latency and throughput.

This on-site role offers a full-time path at a fast-growing AI infrastructure company.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Staff Engineer: GPU Kernels & AI Performance
Staff Engineer: GPU Kernels & AI Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Kernel Engineer
Kernel Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Inference Engineer: High-Performance GPU Systems
Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Performance Engineer: GPU Kernel & Inference Optimize
Performance Engineer: GPU Kernel & Inference Optimize

WORLD LABS • San Francisco (CA)

On-site
USD 200,000 - 300,000
GPU Kernel Engineer for Fast AI Training & Inference
GPU Kernel Engineer for Fast AI Training & Inference

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,900 - 232,800
Annual performance bonus
Equity options
Comprehensive benefits package