Senior AI Cluster Performance Validation Engineer

Advanced Micro Devices

Austin, Northern (TX, KY)

Hybrid

USD 180,000 - 230,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Advanced Micro Devices (AMD) in Austin is seeking a Principal AI Cluster Performance Validation Engineer to optimize GPU clusters and RDMA networks. You will lead performance tuning, profiling, and validation across multi-node systems, collaborating with HW and SW teams.

The role requires strong background in GPU architectures, parallel computing, and hands-on system-level debugging. You’ll shape long-term strategy, drive feature enablement, and stay ahead of industry trends to improve cluster

Qualifications

  • Bachelor's or Master's in CS or EE.
  • Experience optimizing GPU cluster performance.
  • Strong knowledge of RDMA, RoCE, and clustering.
  • Proficiency in Python or Bash for automation.

Responsibilities

  • Scalability testing of GPU clusters under varying workloads and RoCE.
  • Benchmarking and analysis to identify bottlenecks and opportunities for improvement.
  • Cluster network and performance optimization focusing on RDMA throughput, latency, and collective communications.
  • Performance profiling and tuning using established tools and methodologies.
  • Documentation of performance analyses, tuning efforts, and outcomes; reporting to stakeholders.

Skills

GPU architectures
Parallel computing
RDMA networking
Performance tuning
Python
Bash
Profiling tools
Linux networking
Collaboration
ML/HPC design

Education

Bachelor's or Master’s in CS/EE

Job description

Advanced Micro Devices (AMD) in Austin is seeking a Principal AI Cluster Performance Validation Engineer to optimize GPU clusters and RDMA networks. You will lead performance tuning, profiling, and validation across multi-node systems, collaborating with HW and SW teams.

The role requires strong background in GPU architectures, parallel computing, and hands-on system-level debugging. You’ll shape long-term strategy, drive feature enablement, and stay ahead of industry trends to improve cluster

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Performance Validation Engineer
Senior GPU Cluster Performance Validation Engineer

AMD • Austin (TX)

On-site
USD 150,000 - 190,000
AMD Benefits
Principal AI Cluster Performance Validation Engineer
Principal AI Cluster Performance Validation Engineer

AMD • Austin (TX)

On-site
USD 150,000 - 190,000
AMD Benefits
Principal AI Cluster Performance Validation Engineer
Principal AI Cluster Performance Validation Engineer

Advanced Micro Devices • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 230,000
AI Cluster Program Lead: GPU & Rack Validation
AI Cluster Program Lead: GPU & Rack Validation

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 210,000
AMD Benefits
Lead Data Center GPU Performance Architect for AI
Lead Data Center GPU Performance Architect for AI

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 250,000
Senior AI Performance Architect — GPU & Network
Senior AI Performance Architect — GPU & Network

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
GPU Data Center Validation Architect
GPU Data Center Validation Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 190,000
Benefits at a glance
Lead GPU Performance Architect, Datacenter
Lead GPU Performance Architect, Datacenter

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior GPU/AI Systems Engineer - Performance & ML
Senior GPU/AI Systems Engineer - Performance & ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits