Principal Systems Validation Engineer – Data Center GPU System Stress

AMD

Vancouver

On-site

CAD 120,000 - 180,000

Full time

44 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD's DCGPU Validation and Engineering team leads the validation of AI/ML and HPC products, focusing on system stability and stress testing across silicon-to-system boundaries. You will develop validation content, automation, and coverage models to ensure robust performance under maximum stress workloads.

As Principal System Validation Engineer, you will drive strategy, translate it into scalable test plans, and collaborate across engineering groups to tackle complex HW/FW/SW challenges while

Qualifications

  • Bachelor’s or Master’s Degree in Computer Engineering or Electrical Engineering.
  • Experience in GPU/SoC validation, post-silicon bring-up, and data center/platform validation.
  • Proficient in C/C++, scripting languages, and working knowledge of Linux and Windows server environments.

Responsibilities

  • Define, drive, and evolve the system stability and stress validation strategy across test content and automation.
  • Develop deep understanding of silicon SoCs, board-level designs, system interfaces, and firmware/software stacks for stress tests.
  • Collaborate across validation, firmware/software, silicon design, and verification teams to execute robust validation test plans at scale.
  • Provide technical leadership for HW/FW/SW integration challenges and improve test coverage and tooling.
  • Advance validation using AI-assisted workflows and data-driven quality methodologies for faster debugging and readiness.

Skills

C/C++ programming
Python scripting
Perl scripting
Ruby scripting
Linux & Windows environments

Education

Bachelor’s or Master’s Degree in Computer Engineering or Electrical Engineering

Tools

Linux
Windows Server
CI/CD tools

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re redesigning next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger—technology that moves the world forward. Join us and, together, we’ll advance your career.

THE ROLE

Our Data Center GPU (DCGPU) Validation and Engineering team ensures the quality, reliability, security, and performance of AMD's industry-leading AI, ML, and HPC products. Working closely with architecture, design, product engineering, software, and hardware teams, we develop the validation methodologies, automation frameworks, and system-level debug capabilities that enable successful silicon bring-up, product readiness, and exceptional customer quality. Our products demand rigorous testing for optimum performance and reliability under maximum stress. You will play a key role in shaping those strategies and their implementation.

THE PERSON

We are seeking a Principal System Validation Engineer to help drive product-level system and stress testing across AMD's next-generation platforms. In this role you will define and drive the validation strategy for system stability and stress at scale, translate that strategy into test content and automation, and drive collaboration across engineering organizations. You will operate with extensive silicon-to-systems understanding, driving complex system-level integration challenges and ensuring stable system configuration under maximum stress workloads. The Systems Design Engineering team fosters and encourages continuous technical innovation to showcase successes as well as facilitate continuous career development.

KEY RESPONSIBILITIES
  • Define, drive, and evolve AMD's system stability and stress validation strategy, and drive its implementation across test content, coverage models, and automation frameworks to ensure test coverage and configurations lead to full system stability.
  • Develop a deep understanding of our silicon SoCs, board-level designs, system-level designs/interfaces, and firmware/software stacks to drive the development of system stress tests and complex issue investigations.
  • Work across multiple validation, firmware/software, silicon design, and verification teams to develop and execute robust validation test plans at the stress and scale levels that meet our customer requirements.
  • Provide technical leadership for complex system-level challenges spanning HW/FW/SW interactions and apply those learnings back into stronger test coverage, improved tools, and optimized workflows.
  • Drive technical innovation across validation, including design and development of AI-assisted validation workflows and tooling, at-scale validation environments, and data-driven quality methodologies that improve debug efficiency, execution speed, and product readiness.
PREFERRED EXPERIENCE
  • Deep understanding of modern GPU, SoC, and server platform architectures, including experience with post-silicon bring-up, silicon to system-level validation, and AI/ML accelerator or large-scale data center platforms.
  • Extensive knowledge of system validation strategy, with a proven ability to develop validation architectures, methodologies, test strategies, automation frameworks, and end- to-end infrastructure for complex data center and semiconductor product validation
  • Extensive experience with SoC/board/platform-level debug — including delivery, sequencing, analysis, and optimization — with a structured approach to debug workflows.
  • Strong analytical/problem-solving skills, pronounced attention to detail, and a passion for innovation and continuous improvement.
  • History of process improvements and driving early critical-coverage enablement.
  • Hands on programming/scripting/debug skills (e.g., C/C++, Perl, Ruby, Python), along with working knowledge of Linux and Windows server environments.
ACADEMIC CREDENTIALS:
  • Bachelor’s or Master’s Degree in Computer Engineering or Electrical Engineering
LOCATION:
  • Vancouver, BC

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's 'Responsible AI Policy' is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Systems Validation Engineer – Data Center GPU System Stress
Principal Systems Validation Engineer – Data Center GPU System Stress

Advanced Micro Devices • Vancouver

Hybrid
CAD 150,000 - 190,000
Systems Validation Engineer – Data Center GPU
Systems Validation Engineer – Data Center GPU

Advanced Micro Devices • Markham

Hybrid
CAD 90,000 - 130,000
Hybrid work
Benefits
Career growth
Systems Validation Engineer – Data Center GPU
Systems Validation Engineer – Data Center GPU

Socket.dev • Markham

Hybrid
CAD 120,000 - 170,000
Sr. Systems Design Engineer - Data Center GPU
Sr. Systems Design Engineer - Data Center GPU

AMD • Markham

On-site
CAD 110,000 - 170,000
Sr. Systems Design Engineer - Data Center GPU
Sr. Systems Design Engineer - Data Center GPU

Advanced Micro Devices • Markham

Hybrid
CAD 120,000 - 180,000
Benefits at a glance
Hybrid work model
Lead SMU Validation Engineer - Data Center GPU
Lead SMU Validation Engineer - Data Center GPU

AMD • Markham

On-site
CAD 120,000 - 180,000
AMD benefits
Sr. Systems Design Engineer - Data Center GPU
Sr. Systems Design Engineer - Data Center GPU

Socket.dev • Markham

Hybrid
CAD 120,000 - 160,000
SoC Data Path Engineer
SoC Data Path Engineer

Advanced Micro Devices • Markham

Hybrid
CAD 110,000 - 150,000
AMD benefits at a glance
Lead Board Hardware Debug Engineer – Datacenter & AI Platforms
Lead Board Hardware Debug Engineer – Datacenter & AI Platforms

Advanced Micro Devices • Markham

On-site
CAD 100,000 - 130,000
Benefits offered
SoC Data Path Engineer
SoC Data Path Engineer

AMD • Markham

Hybrid
CAD 110,000 - 150,000