ROLE OVERVIEW
As a Senior Failure Analysis Engineer, you will diagnose, isolate, and resolve complex failures across GPU‑accelerated server platforms in rack‑level and data center environments, including ODM factory support. This highly technical, hands‑on role focuses on server bring‑up, system‑level debugging, and rack integration troubleshooting across CPU, GPU, memory, PCIe, networking, power delivery, and thermal subsystems.
KEY RESPONSIBILITIES
- Perform component, system, and rack‑level failure analysis on GPU‑accelerated server platforms
- Debug issues across CPU, GPU, memory, PCIe, networking, storage, power, and thermal subsystems
- Support server bring‑up and manufacturing test failures (POST, BIOS/UEFI configuration, PCIe enumeration, firmware interactions)
- Analyze BIOS, BMC, IPMI, and system logs to identify hardware/firmware interaction issues
- Reproduce factory or field failures in lab environments to validate root cause and corrective actions
- Utilize system‑level debug tools (oscilloscopes, logic analyzers, protocol analyzers, power tools)
- Support ODM manufacturing test, failure debug, and troubleshooting
- Lead issue triage and structured debug activities with ODM partners
- Train ODM teams on debug procedures and failure isolation techniques
- Develop and maintain SOPs, debug guides, and troubleshooting documentation
- Manage factory escalations with clear diagnosis, recommendations, and communication
- Partner with design, firmware, validation, manufacturing, and quality teams to resolve issues
- Drive RCA/FMEA documentation with clear problem statements and data‑backed actions
- Provide feedback to improve DfR, DfT, and DfS
- Use AI tools to:
- Analyze logs and failure patterns
- Improve debug efficiency
- Contribute to scalable debug knowledge bases
QUALIFICATIONS & EXPERIENCE
- Hands‑on experience debugging server platforms in rack‑level or data center environments
- Experience with GPU servers and multi‑node systems
- Strong understanding of power sequencing, PCIe, memory, and thermal systems
- Familiarity with BIOS/UEFI, BMC, IPMI, and firmware‑level debugging
- Proficiency with lab debug tools (oscilloscopes, logic analyzers, protocol analyzers, power analyzers)
- Experience using AI tools for log analysis, debug workflows, or knowledge development
- Strong cross‑functional collaboration and communication skills
- Excellent technical documentation skills
- Willingness to travel up to 25%
- US citizenship required
ACADEMIC CREDENTIALS
- Bachelor’s degree in Electrical Engineering required
- Master’s degree in Electrical or Systems Engineering preferred
LOCATION
Austin, TX (Onsite)
VACANCY STATUS
This role is not eligible for visa sponsorship.
LEGAL STATEMENTS
AMD is an equal opportunity, inclusive employer and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. This posting is for an existing vacancy.