A technology company is looking for a Fellow GPU Performance Optimization Engineer in San Jose, CA. This role focuses on maximizing the performance of large-scale AI training workloads on AMD GPU platforms. Candidates must have deep expertise in GPU architecture, distributed systems, and ML workloads, alongside strong technical leadership skills. The position offers an opportunity to drive innovations across the software-hardware stack and work on impactful optimizations in an inclusive environment.
Qualifications
Deep expertise in GPU performance analysis and optimization.
Strong understanding of GPU architecture, interconnects, and memory hierarchies.
Proficiency in Python and scripting languages.
Responsibilities
Lead performance optimization of AI workloads on AMD GPU platforms.
Identify and eliminate system bottlenecks across compute and memory.
Develop and apply advanced profiling and performance modeling methodologies.
Skills
GPU performance optimization
Distributed systems
Machine Learning workloads
Performance profiling tools
Technical leadership
Education
Ph.D. in Computer Science or Computer Engineering
Tools
PyTorch
TensorFlow
C++
CUDA
Job description
A technology company is looking for a Fellow GPU Performance Optimization Engineer in San Jose, CA. This role focuses on maximizing the performance of large-scale AI training workloads on AMD GPU platforms. Candidates must have deep expertise in GPU architecture, distributed systems, and ML workloads, alongside strong technical leadership skills. The position offers an opportunity to drive innovations across the software-hardware stack and work on impactful optimizations in an inclusive environment.