Senior GPU Inference Engine Engineer

FriendliAI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
Supportive and highly collaborative work environment
Startup equity and health insurance

Job summary

FriendliAI in San Francisco is looking for an experienced Inference Engine Engineer to optimize GPU kernel performance for AI workloads. The position requires a strong background in GPU programming and collaboration with cloud infrastructure teams. The ideal candidate should have 5+ years of experience and expertise in Python, C++. This role offers flexible hours and various benefits, including competitive compensation and health insurance.

Qualifications

  • 5+ years of experience in production or high-impact research environments.
  • Production-level expertise in Python and C++.
  • Experience developing machine learning frameworks or performance-critical runtime systems.
  • Hands-on experience writing and optimizing GPU kernels.
  • Experience working with generative AI models such as transformer and diffusion models.

Responsibilities

  • Design and optimize custom GPU kernels for AI workloads.
  • Contribute to the development of the kernel compiler and other core components.
  • Collaborate with cloud and infrastructure engineers for end-to-end inference performance.
  • Analyze performance bottlenecks and implement optimizations.
  • Maintain production-grade performance infrastructure.

Skills

GPU programming
Python
C++
Machine learning frameworks
Performance optimization

Education

Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

Tools

Machine learning compilers
Profiling tools

Job description

FriendliAI in San Francisco is looking for an experienced Inference Engine Engineer to optimize GPU kernel performance for AI workloads. The position requires a strong background in GPU programming and collaboration with cloud infrastructure teams. The ideal candidate should have 5+ years of experience and expertise in Python, C++. This role offers flexible hours and various benefits, including competitive compensation and health insurance.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Senior AI Kernel Engineer — GPU Inference & Kernel Optimization
Senior AI Kernel Engineer — GPU Inference & Kernel Optimization

Modular • United States

Hybrid
USD 198,000 - 286,000
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Software Engineer – AI Inference Engine
Software Engineer – AI Inference Engine

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Staff Engineer, Mid-Training Infra for Large-Scale AI
Staff Engineer, Mid-Training Infra for Large-Scale AI

Reflection • San Francisco (CA)

On-site
Top-tier compensation
Comprehensive health, dental, and vision insurance
Fully paid parental leave
+2