AI Cluster Program Lead: GPU & Rack Validation

Advanced Micro Devices

Austin (TX)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

AMD Benefits

Job summary

AMD in Austin, TX seeks a hands-on Technical Program Manager to lead AI cluster engineering programs, focusing on GPU platforms, rack bring-up, and validation for hyperscale deployments. You will drive end-to-end delivery from server integration through scale testing, root-cause analysis, and system debug closure, coordinating hardware, firmware, networking, and software teams.

You will collaborate with lab operations, automation, and performance teams to ensure readiness, manage risks, and

Qualifications

  • Experience leading complex hardware or AI infrastructure programs with ownership across bring-up, validation, and deployment phases.
  • Strong technical understanding of GPU-based AI systems, rack architectures, and datacenter infrastructure.
  • Proven ability to manage ambiguity, drive debug execution, and lead cross-functional teams without direct authority.
  • Strong written and verbal communication skills, including executive-level status reporting.

Responsibilities

  • Define, plan, and drive program plans for AI infrastructure systems validation and readiness, including server integration, rack bring-up, and cluster-scale deployment readiness.
  • Create and maintain core PM artifacts: schedules, dependency maps, resource forecasts, risk/issue logs, and program dashboards/status reports.
  • Identify and drive mitigation plans for issues/risks, including cross-team escalations and corrective actions across multiple engineering areas.
  • Drive regular execution reviews with engineering teams and provide concise, data-driven updates to senior leadership.
  • Own program execution for GPU-based AI platforms, spanning system bring-up, qualification, scale readiness, and deployment validation across server, rack, and cluster levels.

Skills

Program management
GPU platforms
System bring-up
Root-cause analysis
Executive reporting
Cross-functional leadership
Jira & Confluence
Data-driven decisions

Education

Bachelor's or Master's in CS/EE/Systems

Tools

Jira
Confluence
Excel/PowerPoint

Job description

AMD in Austin, TX seeks a hands-on Technical Program Manager to lead AI cluster engineering programs, focusing on GPU platforms, rack bring-up, and validation for hyperscale deployments. You will drive end-to-end delivery from server integration through scale testing, root-cause analysis, and system debug closure, coordinating hardware, firmware, networking, and software teams.

You will collaborate with lab operations, automation, and performance teams to ensure readiness, manage risks, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Cluster TPM: Hyperscale AI Infrastructure
AI Cluster TPM: Hyperscale AI Infrastructure

AMD • Austin (TX)

On-site
USD 140,000 - 210,000
AMD benefits
Senior GPU Cluster Performance Validation Engineer
Senior GPU Cluster Performance Validation Engineer

AMD • Austin (TX)

On-site
USD 150,000 - 190,000
AMD Benefits
AI Cluster Technical Program Manager – Validation, Debug & Agentic AI
AI Cluster Technical Program Manager – Validation, Debug & Agentic AI

AMD • Austin (TX)

On-site
USD 140,000 - 210,000
AMD benefits
AI Cluster Technical Program Manager - Validation, Debug & Agentic AI
AI Cluster Technical Program Manager - Validation, Debug & Agentic AI

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 210,000
AMD Benefits
AI/HPC Infrastructure Architect
AI/HPC Infrastructure Architect

AMD • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits at a glance
AI/HPC Data Center Lab Program Lead
AI/HPC Data Center Lab Program Lead

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Global Senior Director, Data Center GPU Programs
Global Senior Director, Data Center GPU Programs

AMD • Austin (TX)

Hybrid
USD 140,000 - 230,000
Data Center GPU Engineer for AI & HPC Deployments
Data Center GPU Engineer for AI & HPC Deployments

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits