Sr. Data Center GPU Validation and Debug Engineer

Advanced Micro Devices

Austin (TX)

Hybrid

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Advanced Micro Devices (AMD) seeks an experienced engineer to validate and debug data center GPU products across the complete software and hardware stack. You will trace unexpected behavior from application to hardware, working across firmware, kernel, compiler, library, and framework teams to drive issues to resolution.

The role emphasizes strong Linux systems engineering, C/C++, Python, and shell automation, with exposure to ROCm/HIP, CUDA, and PCIe. Hybrid work model in the U.S. is supported.

Qualifications

  • Strong Linux and systems-engineering experience.
  • Experience with data center GPUs or high-performance systems.
  • Understanding of GPU architecture and memory hierarchy.
  • Experience debugging across software, firmware, kernel, and hardware boundaries.
  • Ability to analyze logs, traces, and performance counters.

Responsibilities

  • Debug GPU failures across the full software and hardware stack.
  • Reproduce failures and create minimal, actionable test cases.
  • Triage system-level interactions across GPU, CPU, memory, networking, power, and topology.
  • Validate AI, HPC, and communication workloads across platforms and releases.
  • Characterize performance and identify bottlenecks.
  • Analyze single- and multi-GPU behavior under various configurations.
  • Establish benchmarks, baselines, and regression-detection methods.
  • Build automated diagnostics, stress tests, and regression infrastructure.
  • Support customer-representative validation and escalation reproduction.

Skills

Linux systems
C/C++
Python
Shell scripting
GPU architecture
Performance analysis
Debugging

Education

Bachelor’s degree

Tools

ROCm/HIP
CUDA
PCIe

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re redesigning next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger— technology that moves the world forward. Join us and, together, we’ll advance your career.

THE ROLE

We are seeking an experienced engineer to validate and debug data center GPU products across the complete software and hardware stack, spanning workload behavior and performance qualification as well as platform health and root-cause isolation. The role demands strong technical judgment, disciplined failure isolation, and the ability to trace unexpected behavior from application to hardware, working across firmware, kernel, compiler, library, and framework teams to drive issues to resolution.

KEY RESPONSIBILITIES:
  • Debug GPU failures including hangs, crashes, memory faults, firmware errors, and performance regressions across the full software and hardware stack.
  • Reproduce failures and reduce them to minimal, actionable test cases.
  • Triage system-level interactions across GPU, CPU, memory, networking, power, and platform topology.
  • Validate AI, HPC, and communication workloads across GPU platforms, firmware, and software releases.
  • Characterize performance and identify compute, memory, and communication bottlenecks.
  • Analyze single- and multi-GPU behavior across workload configurations, topology, NUMA, power, and thermal conditions.
  • Establish benchmarks, baselines, and regression-detection methods; correlate microbenchmark results with real workload behavior.
  • Build automated diagnostics, stress tests, performance suites, and regression infrastructure.
  • Support customer-representative validation and escalation reproduction.
REQUIRED QUALIFICATIONS:
  • Strong Linux and systems-engineering experience, with C/C++, Python, and shell automation.
  • Experience with data center GPUs, accelerators, or comparable high-performance systems.
  • Understanding of GPU architecture, memory hierarchy, parallel execution, communication, and synchronization.
  • Experience debugging across software, firmware, kernel, and hardware boundaries.
  • Ability to analyze logs, traces, hardware telemetry, and performance counters.
PREFERRED EXPERIENCE:
  • Strong knowledge of Linux kernel drivers, PCIe, firmware interaction, and memory management.
  • Familiarity with GPU programming models and runtimes such as ROCm/HIP, RCCL, or CUDA equivalents.
  • Experience with AI training or inference frameworks such as PyTorch, vLLM, or SGLang.
  • Experience benchmarking AI inference, training, or HPC workloads, including roofline analysis and model-level profiling.
  • Experience with profiling tools, statistical analysis, and regression detection.
  • Familiarity with multi-GPU topology, PCIe/fabric interconnects, NUMA, GPU firmware, RAS, and rack-scale platforms.
  • Experience with system bring-up, qualification, or production data center operations.
  • Familiarity with CI systems, automated validation infrastructure, and fleet-scale testing.
ACADEMIC CREDENTIALS:
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

This role is not eligible for visa sponsorship.

#LI-G11

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Center GPU Validation and Debug Engineer
Sr. Data Center GPU Validation and Debug Engineer

AMD • Austin (TX)

Hybrid
USD 120,000 - 170,000
AMD Benefits at a glance
Lead Systems Debug Engineer - Data Center GPU
Lead Systems Debug Engineer - Data Center GPU

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 220,000
AMD benefits at a glance
Lead Systems Debug Engineer - Data Center GPU
Lead Systems Debug Engineer - Data Center GPU

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
Sr. System/Test Validation Engineer
Sr. System/Test Validation Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
Sr. System/Test Validation Engineer
Sr. System/Test Validation Engineer

Socket.dev • Austin (TX)

On-site
USD 110,000 - 140,000
GPU boards Failure Analysis Engineer
GPU boards Failure Analysis Engineer

Advanced Micro Devices • Secaucus (NJ)

Hybrid
USD 120,000 - 180,000
AMD benefits at a glance
Sr. System/Test Validation Engineer
Sr. System/Test Validation Engineer

AMD • Austin (TX)

On-site
USD 120,000 - 150,000
Platform / System Debug Validation Engineer
Platform / System Debug Validation Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
Senior Datacenter Platform/Debug Engineer
Senior Datacenter Platform/Debug Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 110,000 - 160,000
Platform / System Debug Validation Engineer
Platform / System Debug Validation Engineer

AMD • Austin (TX)

On-site
USD 120,000 - 170,000