GPU Kernel Engineer for AI Inference & Performance
FriendliAI
San Francisco (CA)
On-site
USD 120,000 - 150,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Flexible working hours
Daily lunch and dinner
Health check-up support
Competitive compensation
Startup equity
Health insurance
Job summary
FriendliAI is seeking a GPU Kernel Engineer in San Francisco to design and optimize GPU kernels for AI inference. This role requires expertise in CUDA, C++, and performance-critical systems. You will work on cutting-edge GPU technology and contribute to a highly collaborative, supportive work environment with competitive compensation and startup equity. Join us at the forefront of AI infrastructure innovation.
Qualifications
3+ years of experience in GPU programming.
Strong proficiency in CUDA for NVIDIA GPUs or ROCm/HIP for AMD GPUs.
Deep understanding of GPU architecture and tuning.
Responsibilities
Design, implement, and optimize high-performance GPU kernels for AI inference.
Develop and maintain GPU code in CUDA and C++.
Benchmark and ensure performance parity between NVIDIA and AMD hardware.
Skills
GPU programming
Performance-critical systems
CUDA proficiency
C++ proficiency
Performance tuning
Education
Bachelor’s or Master’s in Computer Science, Computer Engineering, Electrical Engineering
Tools
CUDA
C++
ROCm/HIP
Job description
FriendliAI is seeking a GPU Kernel Engineer in San Francisco to design and optimize GPU kernels for AI inference. This role requires expertise in CUDA, C++, and performance-critical systems. You will work on cutting-edge GPU technology and contribute to a highly collaborative, supportive work environment with competitive compensation and startup equity. Join us at the forefront of AI infrastructure innovation.