Senior AI Performance & Reliability Engineer

AMD

San Jose (CA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking a Principal/Fellow level software engineer for the AI Infrastructure team in San Jose/Bellevue area. You will optimize AI workloads (training and inference) across models, frameworks, and hardware, working with customers to identify bottlenecks and deliver scalable solutions.

The role requires strong software engineering, profiling, and experience with ML frameworks like PyTorch, TensorFlow, or JAX.

Qualifications

  • Profile and optimize AI model training and inference workloads.
  • Improve model throughput, latency, memory efficiency, scalability, and reliability.
  • Identify bottlenecks across models, frameworks, compilers, runtimes, OS, and hardware.
  • Collaborate with customers to understand requirements and reproduce issues.

Responsibilities

  • Develop performance tooling, benchmarks, automation, and observability systems.
  • Translate customer feedback into product and infrastructure improvements.
  • Work with ML engineers, systems engineers, hardware teams, and product teams.

Skills

Python
C++
Performance profiling
ML frameworks

Education

PhD in AI/CS or related field

Tools

CUDA
ROCm/HIP
NCCL
XLA/MLIR

Job description

AMD is seeking a Principal/Fellow level software engineer for the AI Infrastructure team in San Jose/Bellevue area. You will optimize AI workloads (training and inference) across models, frameworks, and hardware, working with customers to identify bottlenecks and deliver scalable solutions.

The role requires strong software engineering, profiling, and experience with ML frameworks like PyTorch, TensorFlow, or JAX.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fellow AI Performance & Reliability Engineer (Hybrid)
Fellow AI Performance & Reliability Engineer (Hybrid)

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Fellow Software Engineer — AI Performance & Reliability
Fellow Software Engineer — AI Performance & Reliability

AMD • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Fellow Software Engineer - AI Performance & Reliability
Fellow Software Engineer - AI Performance & Reliability

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior AI Field Applications Engineer – GPUs & HPC
Senior AI Field Applications Engineer – GPUs & HPC

AMD • Austin (TX)

On-site
USD 120,000 - 160,000
Remote work options
Travel opportunities
AI Performance Software Engineer – GPU & DL Optimizations
AI Performance Software Engineer – GPU & DL Optimizations

AMD • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Inclusive culture
Career advancement opportunities
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
AI Performance Software Engineer — GPU Optimization
AI Performance Software Engineer — GPU Optimization

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 100,000 - 140,000
AMD benefits at a glance
AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 190,000
Senior AI Performance Architect — GPU & Network
Senior AI Performance Architect — GPU & Network

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
Senior GPU/AI Systems Engineer - Performance & ML
Senior GPU/AI Systems Engineer - Performance & ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits