Principal AI Cluster Performance Validation Engineer

Advanced Micro Devices

Austin, Northern (TX, KY)

Hybrid

USD 180,000 - 230,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Advanced Micro Devices (AMD) in Austin is seeking a Principal AI Cluster Performance Validation Engineer to optimize GPU clusters and RDMA networks. You will lead performance tuning, profiling, and validation across multi-node systems, collaborating with HW and SW teams.

The role requires strong background in GPU architectures, parallel computing, and hands-on system-level debugging. You’ll shape long-term strategy, drive feature enablement, and stay ahead of industry trends to improve cluster

Qualifications

  • Bachelor's or Master's in CS or EE.
  • Experience optimizing GPU cluster performance.
  • Strong knowledge of RDMA, RoCE, and clustering.
  • Proficiency in Python or Bash for automation.

Responsibilities

  • Scalability testing of GPU clusters under varying workloads and RoCE.
  • Benchmarking and analysis to identify bottlenecks and opportunities for improvement.
  • Cluster network and performance optimization focusing on RDMA throughput, latency, and collective communications.
  • Performance profiling and tuning using established tools and methodologies.
  • Documentation of performance analyses, tuning efforts, and outcomes; reporting to stakeholders.

Skills

GPU architectures
Parallel computing
RDMA networking
Performance tuning
Python
Bash
Profiling tools
Linux networking
Collaboration
ML/HPC design

Education

Bachelor's or Master’s in CS/EE

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

We are seeking a highly motivated and skilled Principal AI Cluster Performance Validation Engineer to join our dynamic team. In this role, you will be at the forefront of optimizing and achieving peak performance for GPU clusters. The focus of this role is the RDMA networks used in AI Clusters, understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures, parallel clustered computing, and hands-on experience in system level performance tuning and debug methodologies.

THE PERSON:

The team fosters and encourages continuous technical innovation to showcase successes as well as facilitate continuous career development. A seasoned professional who enjoys hands-on problem-solving. In this role, you’ll shape long-term strategy, drive feature enablement and jump in to tackle challenges head-on. You’ll have a direct impact on performance, automation, and validation, while staying ahead of industry trends to provide strategic insights to senior management. The person should be experienced in debugging complex HW/FW and clustered configurations.

KEY RESPONSIBILITIES:
  • Scalability Testing: Evaluate the scalability of GPU clusters by conducting thorough testing under various workloads, ensuring optimal performance across different cluster sizes, configurations, and networking technologies (RoCE)
  • Benchmarking and Analysis: Develop and execute comprehensive benchmarking strategies to assess baseline performance, analyze bottlenecks, and identify areas for improvement within GPU cluster environments.
  • Cluster Network & Performance Optimization: Collaborate with hardware and software teams to enhance the overall performance of GPU clusters, focusing on aspects such as RDMA throughput, RoCE v2 network congestion, latency, and collective communications.
  • Performance Profiling: Utilize profiling tools and methodologies to analyze and identify performance bottlenecks, providing actionable insights for improvement.
  • Performance Tuning: Implement optimization strategies, including but not limited to protocol enhancements, load balancing techniques, and parallel processing optimizations.
  • Documentation: Create detailed documentation of performance analysis, tuning efforts, and outcomes, providing clear and concise reports for internal teams and stakeholders.
  • Collaboration: Work closely with cross-functional teams, including hardware engineers, software developers, and system architects, to integrate performance improvements into the GPU cluster architecture.
  • Continuous Learning: Stay current with the latest developments in GPU architectures, parallel processing, and emerging technologies to drive continuous improvement in GPU cluster performance.
PREFERRED EXPERIENCE:
  • Proven experience in optimizing the performance of GPU clusters.
  • RDMA network configuration, troubleshooting and performance tuning.
  • Strong understanding of GPU architectures, parallel computing concepts, and network protocols.
  • Proficiency in scripting languages (e.g., Python, Bash) for automation and performance analysis.
  • Experience with system level performance analysis tools and methodologies for GPU clusters.
  • Analytical mindset with excellent problem-solving and debug skills.
  • Familiarity with cluster management tools and systems.
  • Excellent communication and collaboration skills for effective teamwork.
  • Linux kernel networking expertise
  • Machine learning and/or HPC system design
ACADEMIC CREDENTIALS:
  • Bachelors or Master’s degree in computer science or electrical engineering
LOCATION:
  • Austin TX, Seattle WA, Santa Clara CA or Secaucus NJ

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Cluster Performance Validation Engineer
Principal AI Cluster Performance Validation Engineer

AMD • Austin (TX)

On-site
USD 150,000 - 190,000
AMD Benefits
Sr Principal AI Performance Software Architect
Sr Principal AI Performance Software Architect

AMD • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Inclusive culture
Career advancement opportunities
AI Cluster Technical Program Manager - Validation, Debug & Agentic AI
AI Cluster Technical Program Manager - Validation, Debug & Agentic AI

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 210,000
AMD Benefits
Principal Datacenter GPU Performance Architect
Principal Datacenter GPU Performance Architect

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
Principal Datacenter GPU Performance Architect
Principal Datacenter GPU Performance Architect

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 150,000 - 190,000
Principal Data Center GPU Performance Architect
Principal Data Center GPU Performance Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 250,000
Fellow GPU Performance Optimization Engineer
Fellow GPU Performance Optimization Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
HPC Systems Engineer - AI Workloads
HPC Systems Engineer - AI Workloads

Advanced Micro Devices • San Jose (CA)

On-site
USD 190,000 - 230,000
Fellow GPU Performance Optimization Engineer
Fellow GPU Performance Optimization Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 180,000
Health insurance
Retirement plan
Paid time off
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance