Staff Product Development Engineer (Debug & Failure Analysis)

ADVANCED MICRO DEVICES (SINGAPORE) PTE LTD

Singapore

On-site

SGD 90,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ADVANCED MICRO DEVICES (SINGAPORE) PTE LTD is seeking a Returns Debug and RMA-focused engineer in Singapore. You will perform system-level and board-level failure analysis on data center CPU/GPU products, collaborating with ASIC teams to isolate root causes.

You will run SLT on various platforms, document findings, and support cross-functional teams to drive product quality improvements. A proactive teammate with strong debugging and scripting skills will thrive in this fast-paced environment.

Qualifications

  • Bachelor's or master's degree in electrical and electronic engineering or computer engineering preferred.
  • Strong debugging and failure analysis knowledge on silicon or board level.
  • Analytical, detail-oriented with fast learning ability.
  • Experience with Linux and modern software tools/benchmarks.

Responsibilities

  • Perform system-level and board-level Failure Analysis on Data center CPU/GPU products including customer returns and field failures.
  • Conduct System-Level Test (SLT) on internal boards or customer platforms to duplicate failures and identify causes.
  • Collaborate with ASIC Engineering teams for die-level investigations and root cause analysis.
  • Investigate excursions and critical issues, supporting DPPM improvement initiatives.
  • Drive test coverage analysis to improve failure detection and mitigation.
  • Assist validation, firmware, and hardware teams to resolve hardware/software issues.
  • Prototype and evaluate new FA tools to enhance GPU failure analysis capabilities.
  • Document findings and corrective actions in clear technical reports.

Skills

Failure Analysis
System-level debugging
Board-level debugging
GPU Architecture
Linux
C++
Python
JTAG
BIOS firmware
Hardware testing
Problem-solving
Report Writing

Education

Electrical and Electronic Engineering / Computer Engineering degree

Tools

Oscilloscopes
Multimeters
Bench Testing
Server hardware

Job description

THE ROLE:

Returns Debug and RMA execution in Quality & Reliability organization, provide supportive functions to the organization to ensure customer quality issues are being addressed, evoking the required actions via failure analysis to improve product quality.

THE PERSON:

You will also need to possess strong verbal and written communication skills, which are essential when working with a global team. A proactive, outstanding teammate who focuses on teamwork, team building, and growing team success.

AMD's environment is fast-paced, results-oriented, and built upon a legion of forward-thinking people with a passion for winning technology!

KEY RESPONSIBILITIES:
  • Perform System-level and board-level Failure Analysis (FA) for Data center CPU/GPU products, covering customer returns and field failures.
  • Conduct System-Level Test (SLT) using internal test board or on customer platform to duplicate customer reported failures and isolate the cause of failure.
  • Perform board power up test and functional test to isolate board or component level failure.
  • Collaborate with ASIC Engineering teams for in-depth GPU ASIC/die-level investigations, fault isolation, and root cause analysis.
  • Investigate Excursion and Critical Issues, supporting DPPM (VF/NFF) improvement initiatives.
  • Drive test coverage analysis and enhancements to improve failure detection and mitigation.
  • Partner with validation, firmware, and hardware teams to resolve hardware, software, and platform issues.
  • Innovate, prototype, and evaluate new FA tools to improve GPU failure analysis capabilities.
  • Knowledgeable in functional test and stress software to enhance debug efficiency.
  • Provide technical assistance, resources, and equipment to support engineering teams in testing and debugging activities.
  • Plan, set up, and install server racks with air and liquid cooling capabilities for advanced test infrastructure.
  • Work closely with program managers and product line quality (PLQ)/customer Interacting teams to align failure analysis report writing with external customer communication.
  • Document debug findings, root cause analysis, and corrective actions in clear, concise technical reports.
  • Serve as the local product owner, responsible for tracking and releasing platform screening programs and BKC revision related to server rack level, board level and OSV programs.
  • Act as the Go-To technical expert for owned products, supporting test program contents, FA methodologies, and customer queries.
  • Proactively identify opportunities for process improvement, code quality enhancements, and hardware coverage expansion.
  • Other duties as assigned by supervisor.
PREFERRED EXPERIENCE:
  • Strong in either silicon or board level debug or Failure Analysis knowledge.
  • Candidate should be analytical and detail-oriented, strongly interested in debugging complex systems, self-starter, and a fast learner
  • Excellent skill in code development, familiarity with Linux and modern software tools/benchmarks and techniques for development.
  • Understanding of GPU or x86 architecture knowledge is much preferred
  • Knowledge or experience in server or data center hardware or platform is a plus.
  • Experience working with power management features such as POST, P-states, etc.
  • Knowledge of industry standards like PCIE, USB, or high bandwidth memory is a strong plus
  • JTAG knowledge is a plus.
  • Strong understanding of BIOS or memory firmware is a plus.
  • Experience programming experience with C++, C#, Python, HTML, or JAVA.
  • Experience with PC HW debugging, including voltage, networking, storage, and thermal control.
  • Experience with building computer systems (desktops, laptops, servers, etc).
  • Experience in server installation, configuration, and maintenance is a plus.
  • Good analytical and problem-solving skills.
  • Strong understanding of hardware debugging, test equipment (multimeters, oscilloscopes, thermal cameras, etc.), and system-level troubleshooting.
  • Proficient in reading electronic schematics and electronic component datasheets.
  • Proficient in AI tool or machine learning knowledge is a plus.
ACADEMIC CREDENTIALS:
  • Bachelor's or master's degree in electrical and Electronic Engineering, Computer Engineering is preferred.
LOCATION:

Singapore

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Product Development Engineer (Failure Analysis)
Senior Product Development Engineer (Failure Analysis)

ADVANCED MICRO DEVICES (SINGAPORE) PTE LTD • Singapore

On-site
SGD 70,000 - 110,000
Senior Product Development Engineer
Senior Product Development Engineer

AMD • Singapore

On-site
SGD 90,000 - 120,000
Sr. Product Development Engineer
Sr. Product Development Engineer

Advanced Micro Devices • Singapore

On-site
SGD 90,000 - 130,000
Staff Failure Analysis Engineer – GPU/Data Center Systems
Staff Failure Analysis Engineer – GPU/Data Center Systems

ADVANCED MICRO DEVICES (SINGAPORE) PTE LTD • Singapore

On-site
SGD 90,000 - 130,000
Senior Quality Engineer
Senior Quality Engineer

Advanced Micro Devices • Singapore

On-site
SGD 90,000 - 150,000
Senior Quality Engineer
Senior Quality Engineer

Advanced Micro Devices (S) Pte Ltd • Singapore

On-site
SGD 90,000 - 150,000
Senior Quality Engineer
Senior Quality Engineer

AMD • Singapore

On-site
SGD 90,000 - 140,000
Senior Hardware Diagnostics & Failure Analysis Engineer
Senior Hardware Diagnostics & Failure Analysis Engineer

Advanced Micro Devices • Singapore

On-site
SGD 90,000 - 150,000
Senior Returns Debug & Reliability Engineer
Senior Returns Debug & Reliability Engineer

AMD Maquinaria • Singapore

On-site
SGD 90,000 - 120,000
Senior Returns Debug & Reliability Engineer
Senior Returns Debug & Reliability Engineer

AMD • Singapore

On-site
SGD 90,000 - 120,000