A cutting-edge AI company in California is looking for a Member of Technical Staff for Kernel/Compiler/Communication. This critical role requires strong expertise in CUDA and GPU optimization, along with 5+ years of experience in performance engineering. The ideal candidate will design high-performance kernels and optimize systems for large GPU clusters, contributing to the next generation of AI solutions. The company offers competitive compensation, equity, and a collaborative work environment.
Qualifications
5+ years of experience in systems, compiler, or performance engineering.
Strong expertise in CUDA or accelerator programming.
Deep understanding of GPU architecture and memory hierarchy.
Experience writing or optimizing high-performance kernels.
Strong background in compilers, runtimes, or code generation.
Responsibilities
Design and implement high-performance kernels for AI workloads.
Optimize compiler and runtime stacks for ML systems.
Improve communication efficiency across large GPU clusters.
Reduce latency and increase throughput for distributed workloads.
Profile and eliminate system bottlenecks across the stack.
Skills
CUDA or accelerator programming
GPU architecture and memory hierarchy
C++
Python
Debugging and profiling skills at system level
Job description
A cutting-edge AI company in California is looking for a Member of Technical Staff for Kernel/Compiler/Communication. This critical role requires strong expertise in CUDA and GPU optimization, along with 5+ years of experience in performance engineering. The ideal candidate will design high-performance kernels and optimize systems for large GPU clusters, contributing to the next generation of AI solutions. The company offers competitive compensation, equity, and a collaborative work environment.