Senior GPU Inference Engine Engineer

FriendliAI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
Supportive and highly collaborative work environment
Startup equity and health insurance

Job summary

FriendliAI in San Francisco is looking for an experienced Inference Engine Engineer to optimize GPU kernel performance for AI workloads. The position requires a strong background in GPU programming and collaboration with cloud infrastructure teams. The ideal candidate should have 5+ years of experience and expertise in Python, C++. This role offers flexible hours and various benefits, including competitive compensation and health insurance.

Qualifications

  • 5+ years of experience in production or high-impact research environments.
  • Production-level expertise in Python and C++.
  • Experience developing machine learning frameworks or performance-critical runtime systems.
  • Hands-on experience writing and optimizing GPU kernels.
  • Experience working with generative AI models such as transformer and diffusion models.

Responsibilities

  • Design and optimize custom GPU kernels for AI workloads.
  • Contribute to the development of the kernel compiler and other core components.
  • Collaborate with cloud and infrastructure engineers for end-to-end inference performance.
  • Analyze performance bottlenecks and implement optimizations.
  • Maintain production-grade performance infrastructure.

Skills

GPU programming
Python
C++
Machine learning frameworks
Performance optimization

Education

Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

Tools

Machine learning compilers
Profiling tools

Job description

FriendliAI in San Francisco is looking for an experienced Inference Engine Engineer to optimize GPU kernel performance for AI workloads. The position requires a strong background in GPU programming and collaboration with cloud infrastructure teams. The ideal candidate should have 5+ years of experience and expertise in Python, C++. This role offers flexible hours and various benefits, including competitive compensation and health insurance.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave for all new parents
+2
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Software Engineer - AI Inference Engine
Software Engineer - AI Inference Engine

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,900 - 232,800
Annual performance bonus
Equity options
Comprehensive benefits package
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000