Fellow AI Performance & Reliability Engineer (Hybrid)

Advanced Micro Devices

San Jose (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD seeks a Fellow Software Engineer to advance AI performance and reliability across training and inference workloads. You will optimize AI systems spanning models, frameworks, and hardware, collaborating with customers and internal teams to improve throughput and stability at scale.

The role demands deep expertise in AI software stacks, performance analysis, and system-level thinking, with strong programming skills in Python and C++. Hybrid work in California or Washington is supported.

Qualifications

  • PhD or master's in AI/ML, CS, or related field.
  • Experience profiling and optimizing AI workloads.
  • Strong foundations in computer architecture and systems performance.

Responsibilities

  • Profile and optimize AI model training and inference workloads.
  • Improve model throughput, latency, memory efficiency, scalability, and reliability.
  • Collaborate with customers to understand requirements and reproduce issues.
  • Document performance findings and best practices.

Skills

Strong software engineering
Profiling & optimizing
Python
C++
PyTorch
TensorFlow
JAX
Debugging
Communication
Customer collaboration

Education

PhD or MS in CS/ML

Tools

CUDA
ROCm
HIP
Triton
XLA
MLIR
NCCL

Job description

AMD seeks a Fellow Software Engineer to advance AI performance and reliability across training and inference workloads. You will optimize AI systems spanning models, frameworks, and hardware, collaborating with customers and internal teams to improve throughput and stability at scale.

The role demands deep expertise in AI software stacks, performance analysis, and system-level thinking, with strong programming skills in Python and C++. Hybrid work in California or Washington is supported.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Performance & Reliability Engineer
Senior AI Performance & Reliability Engineer

AMD • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Fellow Software Engineer - AI Performance & Reliability
Fellow Software Engineer - AI Performance & Reliability

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 240,000
Fellow Software Engineer — AI Performance & Reliability
Fellow Software Engineer — AI Performance & Reliability

AMD • San Jose (CA)

Hybrid
USD 180,000 - 260,000
AI Performance Software Engineer – GPU & DL Optimizations
AI Performance Software Engineer – GPU & DL Optimizations

AMD • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Inclusive culture
Career advancement opportunities
AI Software Performance Fellow
AI Software Performance Fellow

Advanced Micro Devices, Inc. • Bellevue (WA)

On-site
USD 180,000 - 230,000
Comprehensive benefits package
Inclusive workplace culture
AI Systems Engineer: ML Kernels & HPC Acceleration
AI Systems Engineer: ML Kernels & HPC Acceleration

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
AI Software Optimization Fellow — Scale Leader
AI Software Optimization Fellow — Scale Leader

Advanced Micro Devices • Bellevue (WA)

On-site
USD 180,000 - 220,000
Collaborative culture
Comprehensive benefits package
AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 190,000
AI Performance Software Engineer — GPU Optimization
AI Performance Software Engineer — GPU Optimization

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 100,000 - 140,000
AMD benefits at a glance
Senior AI Field Applications Engineer – GPUs & HPC
Senior AI Field Applications Engineer – GPUs & HPC

AMD • Austin (TX)

On-site
USD 120,000 - 160,000
Remote work options
Travel opportunities