GenAI Inference Optimization Lead — GPU Performance

Advanced Micro Devices

San Jose (CA)

Hybrid

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company is looking for a Principal GenAI Inference Optimization Engineer in San Jose, CA. This role will focus on optimizing performance and efficiency of generative AI on AMD GPU platforms. The ideal candidate will have significant expertise in GPU architecture, GenAI optimization techniques, and performance tuning tools. You will work across various layers and collaborate with cross-functional teams to drive impactful optimizations. This position is hybrid and offers a dynamic work environment.

Qualifications

  • Solid understanding of GPU architecture and performance fundamentals.
  • Hands-on experience with techniques for GenAI inference optimization.
  • Experience working on LLM or multimodal inference workloads.

Responsibilities

  • Optimize performance of GenAI inference workloads on AMD GPU platforms.
  • Improve latency, throughput, and cost efficiency for model serving in production.
  • Analyze and resolve bottlenecks across compute and memory systems.

Skills

GPU architecture understanding
GenAI inference optimization
Python
C++/CUDA/HIP
Performance tuning tools
ML frameworks (PyTorch, JAX, TensorFlow)

Education

B.S., M.S. or Ph.D. in Computer Science/Computer Engineering

Tools

Profiling/debugging tools
Inference/serving frameworks (vLLM, SGLang, Triton)

Job description

A leading technology company is looking for a Principal GenAI Inference Optimization Engineer in San Jose, CA. This role will focus on optimizing performance and efficiency of generative AI on AMD GPU platforms. The ideal candidate will have significant expertise in GPU architecture, GenAI optimization techniques, and performance tuning tools. You will work across various layers and collaborate with cross-functional teams to drive impactful optimizations. This position is hybrid and offers a dynamic work environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead GPU AI Inference Performance Engineer
Lead GPU AI Inference Performance Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Fellow GPU Performance Optimizer for AI Training
Fellow GPU Performance Optimizer for AI Training

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 180,000
Health insurance
Retirement plan
Paid time off
Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Senior GPU Performance Engineer for AI Training
Senior GPU Performance Engineer for AI Training

CareerArc • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
Principal AI Performance Engineer, GPU Systems Lead
Principal AI Performance Engineer, GPU Systems Lead

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Senior GPU ML Training Performance Engineer
Senior GPU ML Training Performance Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 130,000 - 170,000
Comprehensive health benefits
Inclusive workplace culture
Opportunities for career advancement
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior GPU Inference Engine Engineer
Senior GPU Inference Engine Engineer

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3