RMA Failure Analysis Engineer

CoreFleet Solutions

San Jose (CA)

On-site

USD 90,000 - 150,000

Full time

29 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

CoreFleet Solutions seeks an experienced RMA Failure Analysis engineer to diagnose and analyze customer-returned GPU servers and enterprise hardware. You will perform root-cause analyses, read schematics, and test across CPUs, GPUs, memory, PCIe, and power delivery, ensuring high-value hardware is analyzed safely.

You'll document results, coordinate with design and quality teams, and maintain rigorous ESD practices while handling production hardware.

Qualifications

  • 4+ years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
  • Strong understanding of GPU server architecture and enterprise server platforms.
  • Experience performing system-, board-, and component-level troubleshooting.
  • Ability to read electrical schematics, block diagrams, and PCB layouts.
  • Hands-on with BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking.
  • Experience with oscilloscopes, DMMs, power analyzers, logic analyzers, and protocol analyzers.
  • Working knowledge of Linux operating systems and command-line troubleshooting.
  • Root cause analysis methodologies and failure isolation techniques.
  • Ability to safely handle sensitive server and GPU hardware with proper ESD practices.

Responsibilities

  • Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, and related hardware assemblies.
  • Troubleshoot at system-, board-, and component-level to identify root causes of hardware failures.
  • Execute functional testing, diagnostics, and debugging using standard lab equipment and server validation tools.
  • Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
  • Document failure analysis findings, corrective actions, and recommendations.
  • Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution.

Skills

Server hardware design
Validation & testing
Debugging
Failure analysis
System engineering
GPU server architecture
Schematics reading
JIRA
Zendesk
Linux CLI

Tools

Oscilloscopes
Digital multimeters
Power analyzers
Logic analyzers
Protocol analyzers

Job description

We are seeking an experienced RMA Failure Analysis for GPU Servers and enterprise server platforms. The engineer will be responsible for diagnosing, troubleshooting, and performing root cause analysis on customer-returned GPU servers, server motherboards, GPU baseboards, and associated hardware subsystems. This ideal candidate will possess strong server architecture knowledge, component-level debugging expertise, and the ability to safely handle and analyze high-value hardware throughout the failure analysis process.

What you'll do
  • Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, GPU baseboards, and related hardware assemblies.
  • Conduct system-level, board-level, and component-level troubleshooting to identify root causes of hardware failures.
  • Execute functional testing, diagnostics, and debug activities using standard lab equipment and server validation tools.
  • Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
  • Analyze failures involving server subsystems including CPUs, GPUs, DIMMs, NICs, SSDs, power supplies, PCIe devices, and cooling/thermal subsystems.
  • Troubleshoot hardware issues related to BIOS, BMC, CPLD, FPGA, PCIe, memory, storage, networking, and power delivery circuits.
  • Perform component-level debugging including capacitors, resistors, fuses, diodes, MOSFETs, voltage regulators, ICs, and other electronic components.
  • Conduct component swapping, isolation testing, and fault reproduction to validate failure mechanisms and root causes.
  • Perform detailed visual and mechanical inspections to identify damaged, missing, misaligned, overheated, or improperly assembled components.
  • Utilize JIRA and Zendesk to track RMA cases, document failure analysis results, manage issue resolution activities, and maintain clear communication across engineering, quality, and customer support teams.
  • Document failure analysis findings, corrective actions, and recommendations to support continuous product quality improvements.
  • Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution and corrective actions.
  • Follow proper ESD and hardware handling procedures while working with customer-returned products, engineering samples, and production hardware.
Required Qualifications
  • 4+ years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
  • Strong understanding of GPU server architecture and enterprise server platforms.
  • Experience performing system-level, board-level, and component-level troubleshooting.
  • Ability to read and interpret electrical schematics, block diagrams, and PCB layouts.
  • Hands-on experience with server technologies including BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking interfaces.
  • Experience using laboratory equipment such as oscilloscopes, digital multimeters (DMM), power analyzers, logic analyzers, and protocol analyzers.
  • Working knowledge of Linux operating systems and command-line troubleshooting.
  • Strong understanding of root cause analysis methodologies and failure isolation techniques.
  • Ability to safely handle sensitive server and GPU hardware while adhering to ESD and hardware handling best practices.
Preferred Qualifications
  • Experience supporting AI, HPC, or GPU-accelerated server platforms.
  • Experience with customer-returned hardware (RMA) failure analysis processes.
  • Knowledge of power delivery architecture, thermal analysis, and signal integrity concepts.
  • Familiarity with manufacturing defects, field failures, and reliability-related investigations.
Critical Requirements
  • Must be capable of independently troubleshooting GPU servers and server hardware down to the component level.
  • Must understand overall server architecture and subsystem interactions before initiating debug activities.
  • Must demonstrate strong analytical and problem-solving skills in hardware failure analysis.
  • Must be comfortable working with customer-returned hardware and managing multiple RMA investigations simultaneously.
  • Must maintain proper hardware handling practices to prevent damage to customer-returned units and engineering samples.
About CoreFleet Solutions

At CoreFleet Solutions, we're building a company focused on delivering exceptional workforce, logistics, and technology services. As a growing startup, every team member has the opportunity to make a meaningful impact and help shape the future of the business.

We partner with organizations to provide staffing solutions, logistics support, and technology deployment services with a commitment to quality, reliability, and customer success.

Why Join CoreFleet?
  • Opportunity to grow with a fast-growing startup
  • Work directly with company leadership
  • Learn new skills across multiple industries
  • Collaborative, supportive, and entrepreneurial culture
  • Make a real impact, your ideas and contributions matter

If you're looking for a place where you can grow your career while helping build something from the ground up, we'd love to hear from you.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RMA Failure Analysis Engineer
RMA Failure Analysis Engineer

Core Fleet Solutions • San Jose (CA)

On-site
USD 120,000 - 170,000
Senior GPU Server Failure Analyst (RMA & Root Cause)
Senior GPU Server Failure Analyst (RMA & Root Cause)

CoreFleet Solutions • San Jose (CA)

On-site
USD 90,000 - 150,000
Hardware Failure Analysis Engineer – Reliability
Hardware Failure Analysis Engineer – Reliability

Maven Ventures • Southaven (MS)

On-site
USD 90,000 - 130,000
Failure Analysis Engineer
Failure Analysis Engineer

PEAK Technical Services Inc. • San Jose (CA)

On-site
USD 90,000 - 130,000
Failure Analysis Planner & On-site Engineer
Failure Analysis Planner & On-site Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 90,000 - 120,000
Staff Failure Analysis Engineer
Staff Failure Analysis Engineer

ZT Systems • Secaucus (NJ), Northern (KY)

Hybrid
USD 93,000 - 136,000
Bonus potential
401(k) retirement savings plan
Tuition reimbursement
+1
Senior Failure Engineer
Senior Failure Engineer

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 180,000
Benefits at a glance
Senior Field Application Engineer - OEM
Senior Field Application Engineer - OEM

NVIDIA • Town of Texas (WI)

On-site
USD 132,000 - 253,000
Equity
Benefits
Data Center Technician L2 - GPU Specialist
Data Center Technician L2 - GPU Specialist

Milestone Technologies, Inc. • Reno (NV), Northern (KY)

Hybrid
USD 90,000 - 140,000
Failure Analysis Engineer - Server Systems Integration
Failure Analysis Engineer - Server Systems Integration

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 170,000