Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Acceler8 Talent in San Francisco seeks a Member of Technical Staff focusing on Kernels & GPU Performance to push the limits of production AI inference across diverse accelerators.
You will implement low-level kernels, analyze memory hierarchies, and collaborate with compiler, ML systems, and runtime teams to optimize latency and throughput.
This on-site role offers a full-time path at a fast-growing AI infrastructure company.
Member of Technical Staff - Kernels & GPU Performance
Employment Type: Full-time
Work Model: On-site
About the Company
A fast-growing AI infrastructure company is building a next-generation compute platform designed for high-performance, efficient AI inference across heterogeneous hardware.
The platform combines large-scale compute infrastructure with an execution layer that partitions AI workloads and maps each stage to the hardware best suited to run it.
The team works with leading AI organizations on technical challenges spanning frontier models, production infrastructure, and emerging accelerator architectures.
About the Role
As a Member of Technical Staff, you will build and optimize the low-level execution primitives that translate accelerator capability into production inference performance.
Rather than optimizing for a single hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks.
Your work will directly impact latency, throughput, hardware utilization, and efficiency across both established and emerging accelerator architectures.
You will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation, while partnering with compiler, ML systems, runtime, and distributed-systems engineers.
What You’ll Do
What We’re Looking For
Nice to Have
Keywords:
GPU Kernels, CUDA, Triton, CUTLASS, Kernel Optimization, GPU Performance, Performance Engineering, Accelerator Programming, AI Inference, Inference Optimization, Low-Level Systems, GPU Architecture, Memory Hierarchy, Memory Coalescing, Shared Memory, Cache Optimization, Occupancy, Latency Hiding, Instruction-Level Parallelism, ILP, Warp, Wavefront, Thread Block, Grid, Profiling, Nsight, Performance Analysis, Throughput Optimization, Latency Optimization, Hardware Utilization, Multi-GPU, Distributed Systems, Heterogeneous Compute, Accelerator Runtime, ML Systems, Compiler Runtime.