AI/HPC Cluster Architect

AMD

Austin (TX)

On-site

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking an experienced AI Systems Engineer to design scalable AI/HPC clusters that meet customer and design compliance requirements. You will review and select compute, storage, networking, and power delivery components to optimize performance and reliability across global deployments.

Collaborating with cross-functional teams, you will deliver cutting-edge infrastructure for AI and high-performance computing workloads, while applying strong rack and cluster design knowledge and expertise

Qualifications

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science or related field.
  • Experience in HPC, AI infrastructure, or data center systems engineering preferred.
  • Strong understanding of rack and cluster design.

Responsibilities

  • System Architecture & Design: Design scalable AI/HPC clusters including compute, storage, and networking.
  • System Architecture & Design: Evaluate and select CPUs, GPUs, accelerators, interconnects, and memory configurations for optimal cluster performance.
  • Network: Design network topologies to maximize overall cluster performance.
  • Network: Understand the network performance needs of different workloads.
  • Network: Understand advantages and performance trade-offs of network topologies for AI/HPC clusters.
  • Storage: Design and optimize storage solutions to maximize AI/HPC cluster performance.
  • Storage: Understand advantages and trade-offs of cluster storage solutions (e.g., Lustre, Ceph).
  • Collaboration: Work across multiple organizations with SMEs to deliver scalable compute infrastructure.
  • Collaboration: Experience in HPC, AI systems/clusters, or data center engineering.
  • Collaboration: Strong understanding of rack and cluster design.
  • Collaboration: Knowledge of GPU/CPU architectures, PCIe, UALink, InfiniBand, and Ethernet networking.
  • Collaboration: Familiarity with AI/ML frameworks and workload characteristics.
  • Collaboration: Excellent problem-solving, communication, and documentation skills.

Skills

AI systems
HPC clusters
Compute architecture
Networking
Power delivery

Education

Bachelor's/Master's in Electrical/Computer Engineering or CS

Tools

Lustre
Ceph
PCIe
InfiniBand
Ethernet networking

Job description

AMD is seeking an experienced AI Systems Engineer to design scalable AI/HPC clusters that meet customer and design compliance requirements. You will review and select compute, storage, networking, and power delivery components to optimize performance and reliability across global deployments.

Collaborating with cross-functional teams, you will deliver cutting-edge infrastructure for AI and high-performance computing workloads, while applying strong rack and cluster design knowledge and expertise

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC Cluster Architect
AI/HPC Cluster Architect

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000
AI/HPC Cluster Architect — Data Center Power & Network
AI/HPC Cluster Architect — Data Center Power & Network

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI/HPC Cluster Architect - Scalable Data Center Design
AI/HPC Cluster Architect - Scalable Data Center Design

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits
Equal opportunity employer
Visa sponsorship not available
AI/HPC Cluster Architect: Power, Network & Systems
AI/HPC Cluster Architect: Power, Network & Systems

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI/HPC Data Center Rack Design Architect
AI/HPC Data Center Rack Design Architect

Socket.dev • Austin (TX)

On-site
USD 120,000 - 190,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits