Fellow Software Development Engineer, GEMM Optimization

AMD

San Jose (CA)

On-site

USD 190,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

AMD is seeking a Fellow Software Development Engineer to architect and develop features for the GEMM and GEMM + X code generator, optimizing performance on current and next- generation GPUs.

You will own deep hardware-aware optimization, work with hardware architects, and innovate algorithms to extract maximum throughput, including attention-aware GEMM, in a highly collaborative environment with customers and internal teams.

Qualifications

  • Strong command of GPU computer architecture and memory hierarchies.
  • Deep expertise in GEMM algorithms and their mapping onto GPU hardware.
  • Strong software engineering fundamentals with track record of high-performance software.
  • Experience optimizing GPU compute kernels, GEMM and attention.
  • Proficiency in C++ and assembly-level GPU programming; Python for tooling/automation.
  • Solid understanding of parallel computing and hardware-software tradeoffs.
  • Clear written and verbal communication with engineers and customers.
  • Experience with GPU compiler backends.
  • Experience with AMD architectures (CDNA/RDNA) or similar competitors.

Responsibilities

  • Analyze AMD GPU hardware specifications in depth and align kernel design with silicon capabilities.
  • Develop support for new ISA and HW/SW optimization features in the code-generator
  • Innovate and implement new algorithms for GEMM and GEMM + X
  • Profile and root-cause GEMM performance bottlenecks across the stack
  • Partner with customers and internal stakeholders to understand workload requirements and deliver performance

Skills

GPU architecture
GEMM algorithms
C++
Assembly
Python
Performance optimization
HW/SW optimization

Education

BS in CS/CE
MS in CS/CE
PhD in CS/CE

Tools

GEMM code-generator

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

AMD is seeking a Fellow Software Development Engineer to architect and develop features in the tensile-lite assembly code generator for General Matrix Multiplication (GEMM) and GEMM + X (including attention). GEMM performance is one of the most critical levers for AI workload efficiency and this role sits at the center of AMD's effort to deliver best-in-class matrix-multiply performance across current and next-generation hardware.

You will own the deep, hardware-aware optimization work that turns AMD's GPU architecture into delivered performance: understanding the hardware specification at the instruction and pipeline level, partnering directly with hardware architects to influence future designs, and innovating algorithms to extract maximum throughput.

THE PERSON:

This role is a strong fit for an engineer who wants to work at the intersection of GPU microarchitecture, low-level code generation, and applied AI performance — someone equally comfortable reading a hardware spec, writing hand-tuned assembly, and explaining a performance tradeoff to a customer or an architect.

KEY RESPONSIBILITIES
  • Analyze AMD GPU hardware specifications in depth and work closely with hardware architects to align kernel design with current silicon capabilities and toinfluence requirements for future generations.
  • Develop support for new ISA and HW/SW optimization features in the code-generator
  • Innovate and implement new algorithms for implementing GEMM and GEMM + X
  • Profile and root-cause GEMM performance bottlenecks across the stack
  • Partner with customers and internal stakeholders to understand real-world workload requirements, reproduce performance issues, and deliver targeted performance
PREFERRED EXPERIENCE:
  • Strong command of GPU computer architecture: compute units, register files, cache/LDS hierarchies, memory bandwidth, matrix cores/WMMA-style instructions, and instruction scheduling/latency hiding.
  • Deep expertise in GEMM algorithms and their mapping onto GPU hardware (tiling, blocking, register/LDS allocation, wave scheduling, memory hierarchy utilization).
  • Strong software engineering fundamentals with a track record of building production-quality, high-performance software.
  • Demonstrated experience optimizing GPU compute kernels, ideally GEMM and attention
  • Proficiency in C++ and assembly-level GPU programming; working knowledge of Python for tooling/automation.
  • Solid understanding of parallel computing, memory hierarchies, and hardware-software performance tradeoffs.
  • Clear written and verbal communication skills; ability to work effectively with hardware architects, customers, and cross-functional engineering teams.
  • Direct experience with any GPU compiler backend
  • Experience with AMD GPU architectures (CDNA/RDNA) or comparable competitive architectures (NVIDIA Hopper/Blackwell, etc.)
  • Contributions to industry standards, publications, or patents related to GPU compute or matrix-multiply optimization
PREFERRED ACADEMIC CREDENTIALS:

BS/MS/PHD in CS/CE or related field with deep relevant experience

LOCATION:

San Jose, CA

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Development Engineer - C++, GPU Math Libraries
Software Development Engineer - C++, GPU Math Libraries

AMD • Austin (TX)

On-site
USD 90,000 - 120,000
GPU Kernel Development Engineer (Fellow)
GPU Kernel Development Engineer (Fellow)

AMD • Austin (TX)

On-site
USD 190,000 - 270,000
Fellow GPU Performance Optimization Engineer
Fellow GPU Performance Optimization Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
Senior Staff Software Development Engineer
Senior Staff Software Development Engineer

AMD • San Jose (CA)

On-site
USD 180,000 - 300,000
Principal Datacenter GPU Performance Architect
Principal Datacenter GPU Performance Architect

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
Principal Software Development Engineer
Principal Software Development Engineer

AMD • Santa Clara (CA)

On-site
USD 190,000 - 260,000
AMD benefits
Fellow, AI Workload Optimization
Fellow, AI Workload Optimization

Advanced Micro Devices, Inc. • Bellevue (WA)

On-site
USD 180,000 - 230,000
Comprehensive benefits package
Inclusive workplace culture
Senior Software Development Engineer
Senior Software Development Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 245,000
GPU Performance Architect
GPU Performance Architect

AMD • Austin (TX)

On-site
USD 180,000 - 280,000
AMD benefits at a glance
Principal Compiler Software Development Engineer - AI/ML
Principal Compiler Software Development Engineer - AI/ML

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 260,000