GenAI Inference Optimization Lead — GPU Performance
Advanced Micro Devices
San Jose (CA)
Hybrid
USD 150,000 - 200,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading technology company is looking for a Principal GenAI Inference Optimization Engineer in San Jose, CA. This role will focus on optimizing performance and efficiency of generative AI on AMD GPU platforms. The ideal candidate will have significant expertise in GPU architecture, GenAI optimization techniques, and performance tuning tools. You will work across various layers and collaborate with cross-functional teams to drive impactful optimizations. This position is hybrid and offers a dynamic work environment.
Qualifications
Solid understanding of GPU architecture and performance fundamentals.
Hands-on experience with techniques for GenAI inference optimization.
Experience working on LLM or multimodal inference workloads.
Responsibilities
Optimize performance of GenAI inference workloads on AMD GPU platforms.
Improve latency, throughput, and cost efficiency for model serving in production.
Analyze and resolve bottlenecks across compute and memory systems.
Skills
GPU architecture understanding
GenAI inference optimization
Python
C++/CUDA/HIP
Performance tuning tools
ML frameworks (PyTorch, JAX, TensorFlow)
Education
B.S., M.S. or Ph.D. in Computer Science/Computer Engineering
A leading technology company is looking for a Principal GenAI Inference Optimization Engineer in San Jose, CA. This role will focus on optimizing performance and efficiency of generative AI on AMD GPU platforms. The ideal candidate will have significant expertise in GPU architecture, GenAI optimization techniques, and performance tuning tools. You will work across various layers and collaborate with cross-functional teams to drive impactful optimizations. This position is hybrid and offers a dynamic work environment.