Kernel Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco seeks a Member of Technical Staff focusing on Kernels & GPU Performance to push the limits of production AI inference across diverse accelerators.

You will implement low-level kernels, analyze memory hierarchies, and collaborate with compiler, ML systems, and runtime teams to optimize latency and throughput.

This on-site role offers a full-time path at a fast-growing AI infrastructure company.

Qualifications

  • Bachelor’s degree in a relevant technical discipline or equivalent practical experience.
  • Strong fundamentals in software engineering and performance.
  • Experience with low-level execution, memory hierarchies, and scheduling.
  • Ability to profile, debug, and optimize latency and throughput.
  • Experience across multiple accelerator architectures.
  • Familiarity with CUDA/Triton/CUTLASS is a plus.

Responsibilities

  • Build and optimize kernels for production AI workloads.
  • Improve latency, throughput, and hardware utilization.
  • Develop execution strategies across multiple accelerator architectures.
  • Optimize memory efficiency, scheduling behavior, and low-level execution characteristics.
  • Analyze accelerator execution models and memory hierarchies to identify bottlenecks.
  • Profile and validate performance across different hardware platforms.
  • Partner with compiler, runtime, ML systems, and distributed-systems engineers on end-to-end performance optimization.
  • Develop optimization approaches that account for architectural differences between accelerators.
  • Help establish performance-engineering standards and best practices across the execution platform.

Skills

Software-engineering fundamentals
Performance-critical systems
Low-level execution
Memory hierarchy understanding
Profiling & debugging

Education

Bachelor’s degree in a relevant technical discipline

Tools

CUDA
Triton
CUTLASS
Accelerator programming models
GPU profiling tools

Job description

Member of Technical Staff - Kernels & GPU Performance


Employment Type: Full-time


Work Model: On-site


About the Company


A fast-growing AI infrastructure company is building a next-generation compute platform designed for high-performance, efficient AI inference across heterogeneous hardware.


The platform combines large-scale compute infrastructure with an execution layer that partitions AI workloads and maps each stage to the hardware best suited to run it.


The team works with leading AI organizations on technical challenges spanning frontier models, production infrastructure, and emerging accelerator architectures.


About the Role


As a Member of Technical Staff, you will build and optimize the low-level execution primitives that translate accelerator capability into production inference performance.


Rather than optimizing for a single hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks.


Your work will directly impact latency, throughput, hardware utilization, and efficiency across both established and emerging accelerator architectures.


You will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation, while partnering with compiler, ML systems, runtime, and distributed-systems engineers.


What You’ll Do



  • Build and optimize kernels for production AI workloads.

  • Improve latency, throughput, and hardware utilization.

  • Develop execution strategies across multiple accelerator architectures.

  • Optimize memory efficiency, scheduling behavior, and low-level execution characteristics.

  • Analyze accelerator execution models and memory hierarchies to identify bottlenecks.

  • Profile and validate performance across different hardware platforms.

  • Partner with compiler, runtime, ML systems, and distributed-systems engineers on end-to-end performance optimization.

  • Develop optimization approaches that account for architectural differences between accelerators.

  • Help establish performance-engineering standards and best practices across the execution platform.

  • Influence how heterogeneous compute hardware is deployed and utilized in production AI infrastructure.


What We’re Looking For



  • Strong software-engineering fundamentals.

  • Experience developing performance-critical systems close to hardware.

  • Strong understanding of low-level execution behavior.

  • Ability to reason about memory hierarchies, compute utilization, scheduling, and performance tradeoffs.

  • Experience profiling, debugging, and optimizing systems for latency and throughput.

  • Bachelor’s degree in a relevant technical discipline or equivalent practical experience.


Nice to Have



  • CUDA

  • Triton

  • CUTLASS

  • Accelerator programming models

  • GPU execution models including warps, wavefronts, blocks, and grids

  • Memory-access optimization and coalescing

  • Shared-memory optimization

  • Cache optimization

  • Occupancy tuning

  • GPU profiling and performance-analysis tools

  • Multi-GPU execution

  • Distributed execution

  • AI inference optimization

  • Heterogeneous accelerator experience


Keywords:


GPU Kernels, CUDA, Triton, CUTLASS, Kernel Optimization, GPU Performance, Performance Engineering, Accelerator Programming, AI Inference, Inference Optimization, Low-Level Systems, GPU Architecture, Memory Hierarchy, Memory Coalescing, Shared Memory, Cache Optimization, Occupancy, Latency Hiding, Instruction-Level Parallelism, ILP, Warp, Wavefront, Thread Block, Grid, Profiling, Nsight, Performance Analysis, Throughput Optimization, Latency Optimization, Hardware Utilization, Multi-GPU, Distributed Systems, Heterogeneous Compute, Accelerator Runtime, ML Systems, Compiler Runtime.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

anyone-ai • United States

On-site
USD 138,000 - 248,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Senior GPU Performance / Kernel Engineer
Senior GPU Performance / Kernel Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
Dental insurance
Vision insurance
+2
Senior Kernel Engineer: 26-02794
Senior Kernel Engineer: 26-02794

Akraya, Inc. • Bellevue (WA)

On-site
USD 117,000 - 124,000
NPU Kernel/Operator Engineer
NPU Kernel/Operator Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 120,000 - 160,000