Principal AI Performance & Reliability Engineer

AMD

San Jose (CA)

Hybrid

USD 180,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

AMD is seeking a Principal or Fellow level software engineer for the AI Infrastructure team in San Jose, CA or Bellevue, WA. You will optimize AI workloads across model training and inference, collaborating with customers and internal teams to identify bottlenecks and improve performance and reliability.

This role spans AI software and hardware stacks, requiring deep knowledge of architectures, ML frameworks, and production systems.

Qualifications

  • PhD or master's degree in AI/ML/CS or related field.
  • Strong expertise in AI infrastructure performance.
  • Experience with production ML workloads.
  • Excellent collaboration with customers and teams.

Responsibilities

  • Profile and optimize AI training and inference workloads.
  • Improve throughput, latency, memory efficiency and scalability.
  • Identify bottlenecks across models, frameworks, and hardware.
  • Develop performance tooling, benchmarks, and observability.
  • Collaborate with ML engineers, systems engineers, and product teams.
  • Document findings and recommendations.

Skills

Python
C++
PyTorch
TensorFlow
JAX
Profiling
Distributed systems
Computer architecture
Customer collaboration

Education

PhD or master's in AI/CS or related

Tools

ROCm
HIP
CUDA
Triton
XLA
MLIR
NCCL

Job description

AMD is seeking a Principal or Fellow level software engineer for the AI Infrastructure team in San Jose, CA or Bellevue, WA. You will optimize AI workloads across model training and inference, collaborating with customers and internal teams to identify bottlenecks and improve performance and reliability.

This role spans AI software and hardware stacks, requiring deep knowledge of architectures, ML frameworks, and production systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance & Reliability Engineer
Senior AI Performance & Reliability Engineer

AMD • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Principal AI Performance & Reliability Engineer
Principal AI Performance & Reliability Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

Hybrid
USD 250,000 - 420,000
Hybrid work model
Benefits package
Fellow AI Performance & Reliability Engineer (Hybrid)
Fellow AI Performance & Reliability Engineer (Hybrid)

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Principal Software Engineer — AI Performance & Reliability
Principal Software Engineer — AI Performance & Reliability

AMD • San Jose (CA)

Hybrid
USD 180,000 - 300,000
Fellow Software Engineer - AI Performance & Reliability
Fellow Software Engineer - AI Performance & Reliability

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 240,000
Fellow Software Engineer — AI Performance & Reliability
Fellow Software Engineer — AI Performance & Reliability

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
Principal Software Engineer — AI Performance & Reliability
Principal Software Engineer — AI Performance & Reliability

Advanced Micro Devices, Inc. • San Jose (CA)

Hybrid
USD 250,000 - 420,000
Hybrid work model
Benefits package
AI Performance Software Engineer – GPU & DL Optimizations
AI Performance Software Engineer – GPU & DL Optimizations

AMD • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Inclusive culture
Career advancement opportunities
Senior AI Engineer — Autonomous Systems & Inference
Senior AI Engineer — Autonomous Systems & Inference

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior AI Performance Architect — GPU & Network
Senior AI Performance Architect — GPU & Network

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package