Senior AI Cluster Hardware Engineer

AMD

Austin (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD in Austin, TX, is seeking a talented GPU Cluster Network Performance Attainment Engineer to optimize GPU clusters for peak performance. The ideal candidate will evaluate scalability, perform benchmarking, and implement performance tuning strategies while collaborating with hardware and software teams.

This position requires a deep understanding of GPU architectures, RDMA networks, and performance optimization methodologies. A degree in electrical or computer engineering is preferred. Benefits include comprehensive AMD perks, and this role does not offer visa sponsorship.

Qualifications

  • Proven experience in optimizing the performance of GPU clusters.
  • Strong understanding of GPU architectures and parallel computing concepts.
  • Experience with network protocols used in GPU clusters.

Responsibilities

  • Evaluate scalability of GPU clusters under various workloads.
  • Utilize profiling tools to analyze performance bottlenecks.
  • Implement optimization strategies for performance tuning.
  • Collaborate with cross-functional teams to enhance GPU cluster performance.
  • Develop benchmarking strategies to assess performance.

Skills

Optimization of GPU clusters
Understanding of RDMA network drivers
Scripting languages (Python, Bash)
System-level performance analysis tools
Excellent communication and collaboration skills
Linux kernel networking expertise
Analytical mindset and problem-solving skills
Machine learning and/or HPC system design

Education

Bachelor’s or Master’s degree in electrical or computer engineering

Job description

The Role

We are seeking a highly motivated and skilled GPU Cluster Network Performance Attainment Engineer to join our dynamic team. In this role, you will be at the forefront of optimizing and achieving peak performance for GPU clusters, focusing on the RDMA networks used in AI Clusters and understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures, parallel computing, and hands‑on experience in system‑level performance tuning and debug methodologies.

Key Responsibilities
  • Scalability Testing: Evaluate the scalability of GPU clusters by conducting thorough testing under various workloads, ensuring optimal performance across different cluster sizes, configurations, and networking technologies (RoCE & IB).
  • Performance Profiling: Utilize profiling tools and methodologies to analyze and identify performance bottlenecks, providing actionable insights for improvement.
  • Performance Tuning: Implement optimization strategies, including protocol enhancements, load balancing techniques, and parallel processing optimizations.
  • Documentation: Create detailed documentation of performance analysis, tuning efforts, and outcomes, providing clear and concise reports for internal teams and stakeholders.
  • Collaboration: Work closely with cross‑functional teams, including hardware engineers, software developers, and system architects, to integrate performance improvements into the GPU cluster architecture.
  • NIC & Performance Optimization: Collaborate with hardware and software teams to enhance the overall performance of GPU clusters, focusing on RDMA throughput, latency, and collective communications.
  • Benchmarking and Analysis: Develop and execute comprehensive benchmarking strategies to assess baseline performance, analyze bottlenecks, and identify areas for improvement within GPU cluster environments.
  • Continuous Learning: Stay current with the latest developments in GPU architectures, parallel processing, and emerging technologies to drive continuous improvement in GPU cluster performance.
Preferred Experience
  • Proven experience in optimizing the performance of GPU clusters.
  • Understanding of RDMA network drivers.
  • Strong understanding of GPU architectures, parallel computing concepts, and network protocols.
  • Proficiency in scripting languages (e.g., Python, Bash) for automation and performance analysis.
  • Experience with system‑level performance analysis tools and methodologies for GPU clusters.
  • Analytical mindset with excellent problem‑solving and debug skills.
  • Familiarity with cluster management tools and systems.
  • Excellent communication and collaboration skills for effective teamwork.
  • RDMA network configuration, troubleshooting and performance tuning.
  • Linux kernel networking expertise.
  • Machine learning and/or HPC system design.
Academic Credentials
  • Bachelor’s or Master’s degree in electrical or computer engineering preferred.

Location: Austin, TX

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Cluster Hardware Engineer
Senior AI Cluster Hardware Engineer

Socket.dev • Austin (TX)

Hybrid
USD 120,000 - 160,000
AMD benefits at a glance
Senior AI Cluster Hardware Engineer
Senior AI Cluster Hardware Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 100,000 - 130,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI Cluster & Data Center Design Engr
AI Cluster & Data Center Design Engr

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 190,000
AMD benefits at a glance
AI Cluster & Data Center Design Engr
AI Cluster & Data Center Design Engr

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Cluster Technical Program Manager - Validation, Debug & Agentic AI
AI Cluster Technical Program Manager - Validation, Debug & Agentic AI

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 210,000
AMD Benefits
Data Center GPU Performance Attainment Lead
Data Center GPU Performance Attainment Lead

Advanced Micro Devices • Austin (TX)

Hybrid
USD 90,000 - 120,000