Senior GPU Server Failure Analyst (RMA & Root Cause)

CoreFleet Solutions

San Jose (CA)

On-site

USD 90,000 - 150,000

Full time

23 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CoreFleet Solutions seeks an experienced RMA Failure Analysis engineer to diagnose and analyze customer-returned GPU servers and enterprise hardware. You will perform root-cause analyses, read schematics, and test across CPUs, GPUs, memory, PCIe, and power delivery, ensuring high-value hardware is analyzed safely.

You'll document results, coordinate with design and quality teams, and maintain rigorous ESD practices while handling production hardware.

Qualifications

  • 4+ years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
  • Strong understanding of GPU server architecture and enterprise server platforms.
  • Experience performing system-, board-, and component-level troubleshooting.
  • Ability to read electrical schematics, block diagrams, and PCB layouts.
  • Hands-on with BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking.
  • Experience with oscilloscopes, DMMs, power analyzers, logic analyzers, and protocol analyzers.
  • Working knowledge of Linux operating systems and command-line troubleshooting.
  • Root cause analysis methodologies and failure isolation techniques.
  • Ability to safely handle sensitive server and GPU hardware with proper ESD practices.

Responsibilities

  • Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, and related hardware assemblies.
  • Troubleshoot at system-, board-, and component-level to identify root causes of hardware failures.
  • Execute functional testing, diagnostics, and debugging using standard lab equipment and server validation tools.
  • Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
  • Document failure analysis findings, corrective actions, and recommendations.
  • Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution.

Skills

Server hardware design
Validation & testing
Debugging
Failure analysis
System engineering
GPU server architecture
Schematics reading
JIRA
Zendesk
Linux CLI

Tools

Oscilloscopes
Digital multimeters
Power analyzers
Logic analyzers
Protocol analyzers

Job description

CoreFleet Solutions seeks an experienced RMA Failure Analysis engineer to diagnose and analyze customer-returned GPU servers and enterprise hardware. You will perform root-cause analyses, read schematics, and test across CPUs, GPUs, memory, PCIe, and power delivery, ensuring high-value hardware is analyzed safely.

You'll document results, coordinate with design and quality teams, and maintain rigorous ESD practices while handling production hardware.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RMA Failure Analysis Engineer
RMA Failure Analysis Engineer

Core Fleet Solutions • San Jose (CA)

On-site
USD 120,000 - 170,000
RMA Failure Analysis Engineer
RMA Failure Analysis Engineer

CoreFleet Solutions • San Jose (CA)

On-site
USD 90,000 - 150,000
Senior Failure Analysis Engineer, Data Center GPUs
Senior Failure Analysis Engineer, Data Center GPUs

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 180,000
Benefits at a glance
Senior GPU Fleet Engineer — Hardware & Failure Analysis
Senior GPU Fleet Engineer — Hardware & Failure Analysis

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
GPU Board Debug & Failure Analysis Engineer
GPU Board Debug & Failure Analysis Engineer

AMD • Secaucus (NJ)

On-site
USD 120,000 - 150,000
GPU PCBA Failure & Debug Engineer
GPU PCBA Failure & Debug Engineer

Advanced Micro Devices • Secaucus (NJ)

Hybrid
USD 120,000 - 180,000
AMD benefits at a glance
Lead Server Reliability & Failure Analysis Engineer
Lead Server Reliability & Failure Analysis Engineer

ZT Group Intl, Inc. dba ZT Systems • Secaucus (NJ)

On-site
USD 93,000 - 136,000
Competitive base salary
Bonus eligibility
401(k) retirement savings plan
+3
Power & Thermal Failure Analysis Engineer for GPUs
Power & Thermal Failure Analysis Engineer for GPUs

AMD • Secaucus (NJ)

On-site
USD 90,000 - 120,000
Semiconductor Failure Analysis Specialist
Semiconductor Failure Analysis Specialist

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 60,000 - 101,000
Equity
Benefits package
Firmware-Savvy Failure Analysis Engineer - Server Platforms
Firmware-Savvy Failure Analysis Engineer - Server Platforms

AMD • Secaucus (NJ)

On-site
USD 130,000 - 190,000
Benefits at a glance