Failure Analysis Engineer - Server Systems Integration

AMD

Secaucus (NJ)

On-site

USD 130,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Benefits at a glance

Job summary

AMD is seeking a hands-on systems integrator and failure analysis leader for Helios server platform bring-up. You will own end-to-end failure analysis, develop structured debug plans, review logs and telemetry, and guide corrective actions across BIOS, BMC, CPLD, silicon, and manufacturing teams.

The role requires strong RCA skills, experience with test equipment, and the ability to drive cross-functional debugging and documentation in a fast-paced compute hardware environment.

Qualifications

  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related field.
  • Hands-on server platform bring-up and failure analysis experience.
  • Firmware-aware hardware failure analysis expertise.
  • Ability to lead cross-functional RCA reviews.
  • Experience with diagnostics, logs, and test equipment.

Responsibilities

  • Lead end-to-end failure analysis and root cause ownership for Helios server platform bring-up.
  • Own platform bring-up debug across BIOS, BMC, CPLD, PMBus, POST, PCIe.
  • Develop structured debug plans using logs, traces, and telemetry.
  • Automate data collection and analysis with Python and Linux tools.
  • Create clear technical reports and 8D-style documentation.
  • Interface with silicon, firmware, validation, and manufacturing teams to drive corrective actions.
  • Define diagnostic strategy and instrumentation needs.

Skills

Firmware awareness
Structured debugging
Root-cause analysis
Python scripting
Cross-functional leadership
Oscilloscopes & analyzers

Education

Bachelor's or Master's in EE/CE/CS
Equivalent hands-on experience

Tools

Oscilloscopes
Logic analyzers
Protocol analyzers
Register analysis tools
Schematic review tools

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

Own end-to-end FA across hardware, firmware, silicon, and integration for Helios server platform bring-up and system-level failures. This role drives firmware-aware hardware debug across BIOS, BMC, CPLD, PMBus, POST, PCIe, high-speed interconnect initialization, power sequencing, register state, and platform-level interactions. The engineer will determine whether failures are driven by firmware behavior, hardware design, silicon behavior, component quality, power delivery, configuration, or cross-domain interaction issues. Success in this role requires serving as the firmware SME for FA, enabling independent isolation of systemic system issues while owning interfaces with BIOS, BMC, silicon, validation, design, manufacturing, supplier quality, and customer-facing teams to drive corrective actions and improve platform quality.

THE PERSON:

The ideal candidate is a hands-on systems integrator and failure analysis technical leader with deep platform bring-up experience on complex server or hyperscale systems. They can navigate ambiguous failures across BIOS, BMC, CPLD, firmware, silicon, motherboard, PDB, power delivery, PCIe, high-speed interconnects, diagnostics, and system configuration boundaries. This person is not expected to be a firmware coder; instead, they must understand firmware-controlled hardware behavior well enough to isolate whether the failure is firmware, hardware, or an interaction between domains. They should be comfortable leading structured debug, reviewing register dumps and logs, validating behavior with scopes and logic analyzers, defining diagnostic strategy, communicating clear RCA conclusions, and driving corrective actions across cross-functional engineering teams.

KEY RESPONSIBILITIES:
  • Lead end-to-end failure analysis and root cause ownership for Helios server platform bring-up, factory, customer, and system-level failures.
  • Own platform bring-up debug across BIOS, BMC, CPLD, PMBus, POST, PCIe enumeration, high-speed interconnect initialization, resets, clocks, power sequencing, and system configuration domains.
  • Determine whether failures are caused by firmware behavior, hardware design, component quality, silicon behavior, power delivery, manufacturing process, configuration, or cross-domain interaction issues.
  • Develop and execute structured debug plans using register dumps, firmware and BIOS logs, BMC event logs, telemetry, POST codes, diagnostic results, schematics, board layouts, oscilloscope captures, and logic analyzer traces.
  • Perform power sequencing and platform readiness debug, including rail enable timing, reset behavior, clock availability, PMBus communication, voltage/current telemetry, and fault propagation analysis.
  • Validate firmware-controlled hardware behavior using scopes, logic analyzers, protocol tools, register reads, and data-driven correlation across boot, initialization, and failure states.
  • Own technical interfaces with silicon, BIOS, BMC, firmware, validation, diagnostics, hardware design, manufacturing, and supplier teams to lead cross-functional RCA and drive corrective action closure.
  • Automate debug data collection, log parsing, register analysis, and failure correlation using Python, Linux tools, scripting, and data analysis workflows.
  • Create clear technical reports, executive summaries, debug timelines, and 8D-style documentation that communicate failure mode, evidence, root cause, impact, and recommended actions.
  • Define diagnostic strategy and instrumentation needs by identifying gaps in bring-up procedures, telemetry, register visibility, platform logs, factory screens, and customer debug processes to accelerate systemic issue isolation.
PREFERRED SKILLS AND EXPERIENCE:
  • Strong experience debugging server, hyperscale, GPU/accelerator, rack-scale, or comparable complex compute platforms through bring-up, validation, factory, or customer failure analysis phases.
  • Deep platform bring-up competency across BIOS, BMC, CPLD, PMBus, POST, PCIe enumeration, resets, clocks, power sequencing, register state, and high-speed interconnect initialization flows.
  • Ability to isolate failures across firmware-controlled hardware behavior, silicon, motherboard/PDB design, power delivery, component quality, diagnostics, manufacturing process, and system configuration boundaries.
  • Hands-on proficiency with oscilloscopes, logic analyzers, protocol analyzers, digital multimeters, power supplies, telemetry tools, register access utilities, and lab validation equipment.
  • Experience analyzing register dumps, BIOS/BMC/FW logs, POST codes, event logs, telemetry streams, diagnostic outputs, schematic evidence, and electrical captures to build evidence-based RCA conclusions.
  • Working knowledge of PCIe, high-speed interfaces, interconnect initialization, retimers, link training, signal path dependencies, and failure modes common to dense server platforms.
  • Experience using Python, shell scripting, Linux tools, SQL or data analysis methods to automate log parsing, register review, debug triage, and failure correlation.
  • Ability to own technical interfaces and lead cross-functional RCA reviews with silicon, BIOS, BMC, firmware, validation, diagnostics, design, manufacturing and supplier teams, driving corrective actions to closure.
  • Experience with structured problem solving, 8D, platform bring-up debug, validation escapes, factory issue resolution, supplier quality, or customer failure analysis processes.
  • Ability to understand firmware behavior and hardware control flows without being a pure firmware developer; this role requires firmware-aware systems integration and failure analysis expertise.
ACADEMIC CREDENTIALS:
  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline.
  • Equivalent hands-on experience in server platform bring-up, systems integration debug, firmware-aware hardware failure analysis, validation, or complex system-level RCA will be considered.

#LI-LB1

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Failure Analysis Engineer - Server Systems Integration
Failure Analysis Engineer - Server Systems Integration

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 170,000
Failure Analysis Engineer - Server PCBA Design - Power
Failure Analysis Engineer - Server PCBA Design - Power

AMD • Secaucus (NJ)

On-site
USD 130,000 - 210,000
Failure Analysis Engineer
Failure Analysis Engineer

AMD • Secaucus (NJ)

On-site
USD 120,000 - 160,000
Benefits at a glance
Senior Failure Engineer
Senior Failure Engineer

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 180,000
Benefits at a glance
Senior Failure Engineer
Senior Failure Engineer

AMD • Secaucus (NJ)

On-site
USD 120,000 - 180,000
Customer Debug Engineer
Customer Debug Engineer

Advanced Micro Devices • Seattle (WA)

On-site
USD 110,000 - 150,000
Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug
Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug

Advanced Micro Devices, Inc. • Secaucus (NJ)

On-site
USD 130,000 - 160,000
Competitive benefits
Customer Debug Engineer
Customer Debug Engineer

AMD • Seattle (WA)

On-site
USD 120,000 - 150,000
Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug
Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug

AMD • Secaucus (NJ)

On-site
USD 130,000 - 160,000
Comprehensive benefits package
Opportunities for professional growth
Lead System Debug Engineer - Server Validation
Lead System Debug Engineer - Server Validation

AMD • Austin (TX)

On-site
USD 180,000 - 230,000