A leading semiconductor company is seeking a Fellow GPU Performance Optimization Engineer to maximize performance of AI workloads on AMD GPU platforms. The role requires expertise in GPU performance analysis and distributed systems. Candidates should have experience optimizing large-scale training and strong knowledge of GPU architecture. This position is hybrid in San Jose, California, and is not eligible for visa sponsorship.
Qualifications
Deep expertise in GPU performance optimization and distributed training.
Proven experience optimizing workloads across thousands of GPUs.
Strong understanding of ML frameworks for performance tuning.
Responsibilities
Lead performance optimization of large-scale AI training workloads.
Identify and eliminate system bottlenecks for compute and memory.
Optimize distributed training strategies for efficiency.
Skills
GPU architecture knowledge
Performance optimization
Distributed training expertise
Communication patterns understanding
Education
Ph.D. in Computer Science or related field
Tools
ROCm tools
NVIDIA Nsight
PyTorch
TensorFlow
Job description
A leading semiconductor company is seeking a Fellow GPU Performance Optimization Engineer to maximize performance of AI workloads on AMD GPU platforms. The role requires expertise in GPU performance analysis and distributed systems. Candidates should have experience optimizing large-scale training and strong knowledge of GPU architecture. This position is hybrid in San Jose, California, and is not eligible for visa sponsorship.