RMA Failure Analysis Engineer

CoreFleet Solutions

San Jose (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Growth opportunities
Direct leadership interaction
Skill development
Collaborative culture
Impactful work

Job summary

CoreFleet Solutions is seeking an experienced RMA Failure Analysis engineer for GPU servers and enterprise server platforms. You will diagnose, troubleshoot, and perform root-cause analysis on customer-returned GPU servers, server motherboards, GPU boards, and related hardware subsystems.

The role requires strong server architecture knowledge, component-level debugging skills, and the ability to safely handle high-value hardware while conducting failure analysis.

Qualifications

  • 4 years of experience in server hardware design, validation, testing, failure analysis, or system engineering.
  • Strong understanding of GPU server architectures.
  • Experience reading electrical schematics, block diagrams, and PCB layouts.
  • Hands-on experience with lab equipment and troubleshooting across BIOS, BMC, PCIe, memory, storage, and networking.

Responsibilities

  • Perform failure analysis on customer-returned GPU servers and related hardware.
  • Conduct system-, board-, and component-level troubleshooting to identify root causes.
  • Execute functional testing and diagnostics using standard lab equipment.
  • Read schematics, block diagrams, layouts, and manufacturing documentation.
  • Document failure analysis findings and communicate with engineering, quality, and customer support teams.
  • Follow proper ESD and hardware handling procedures.

Skills

Server hardware design
Failure analysis
System engineering
Root cause analysis
Linux troubleshooting

Tools

Oscilloscopes
Digital multimeters
Power analyzers
Logic analyzers
Protocol analyzers

Job description

About the role

We are seeking an experienced RMA Failure Analysis for GPU Servers and enterprise server platforms. The engineer will be responsible for diagnosing, troubleshooting, and performing root cause analysis on customer-returned GPU servers, server motherboards, GPU baseboards, and associated hardware subsystems.

This ideal candidate will possess strong server architecture knowledge, component-level debugging expertise, and the ability to safely handle and analyze high-value hardware throughout the failure analysis process.

What you'll do
  • Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, GPU baseboards, and related hardware assemblies.
  • Conduct system-level, board-level, and component-level troubleshooting to identify root causes of hardware failures.
  • Execute functional testing, diagnostics, and debug activities using standard lab equipment and server validation tools.
  • Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
  • Analyze failures involving server subsystems including CPUs, GPUs, DIMMs, NICs, SSDs, power supplies, PCIe devices, and cooling/thermal subsystems.
  • Troubleshoot hardware issues related to BIOS, BMC, CPLD, FPGA, PCIe, memory, storage, networking, and power delivery circuits.
  • Perform component-level debugging including capacitors, resistors, fuses, diodes, MOSFETs, voltage regulators, ICs, and other electronic components.
  • Conduct component swapping, isolation testing, and fault reproduction to validate failure mechanisms and root causes.
  • Perform detailed visual and mechanical inspections to identify damaged, missing, misaligned, overheated, or improperly assembled components.
  • Utilize JIRA and Zendesk to track RMA cases, document failure analysis results, manage issue resolution activities, and maintain clear communication across engineering, quality, and customer support teams.
  • Document failure analysis findings, corrective actions, and recommendations to support continuous product quality improvements.
  • Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution and corrective actions.
  • Follow proper ESD and hardware handling procedures while working with customer-returned products, engineering samples, and production hardware.
Qualifications
Required Qualifications
  • 4 years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
  • Strong understanding of GPU server architecture and enterprise server platforms.
  • Experience performing system-level, board-level, and component-level troubleshooting.
  • Ability to read and interpret electrical schematics, block diagrams, and PCB layouts.
  • Hands-on experience with server technologies including BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking interfaces.
  • Experience using laboratory equipment such as oscilloscopes, digital multimeters (DMM), power analyzers, logic analyzers, and protocol analyzers.
  • Working knowledge of Linux operating systems and command-line troubleshooting.
  • Strong understanding of root cause analysis methodologies and failure isolation techniques.
  • Ability to safely handle sensitive server and GPU hardware while adhering to ESD and hardware handling best practices.
Preferred Qualifications
  • Experience supporting AI, HPC, or GPU-accelerated server platforms.
  • Experience with customer-returned hardware (RMA) failure analysis processes.
  • Knowledge of power delivery architecture, thermal analysis, and signal integrity concepts.
  • Familiarity with manufacturing defects, field failures, and reliability-related investigations.
Critical Requirements
  • Must be capable of independently troubleshooting GPU servers and server hardware down to the component level.
  • Must understand overall server architecture and subsystem interactions before initiating debug activities.
  • Must demonstrate strong analytical and problem-solving skills in hardware failure analysis.
  • Must be comfortable working with customer-returned hardware and managing multiple RMA investigations simultaneously.
  • Must maintain proper hardware handling practices to prevent damage to customer-returned units and engineering samples.
About CoreFleet Solutions

At CoreFleet Solutions, we're building a company focused on delivering exceptional workforce, logistics, and technology services. As a growing startup, every team member has the opportunity to make a meaningful impact and help shape the future of the business.

We partner with organizations to provide staffing solutions, logistics support, and technology deployment services with a commitment to quality, reliability, and customer success.

Why Join CoreFleet
  • Opportunity to grow with a fast-growing startup
  • Work directly with company leadership
  • Learn new skills across multiple industries
  • Collaborative, supportive, and entrepreneurial culture
  • Make a real impact, your ideas and contributions matter

If you're looking for a place where you can grow your career while helping build something from the ground up, we'd love to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RMA Failure Analysis Engineer: GPU Server Systems
RMA Failure Analysis Engineer: GPU Server Systems

CoreFleet Solutions • San Jose (CA)

On-site
USD 120,000 - 180,000
Growth opportunities
Direct leadership interaction
Skill development
+2
System Failure Analysis Engineer (GPU Servers / Data Center)
System Failure Analysis Engineer (GPU Servers / Data Center)

AMD • Austin (TX)

On-site
USD 100,000 - 130,000
System Failure Analysis Engineer (GPU Servers / Data Center)
System Failure Analysis Engineer (GPU Servers / Data Center)

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 120,000 - 150,000
System Failure Analysis Engineer (GPU Servers / Data Center)
System Failure Analysis Engineer (GPU Servers / Data Center)

Advanced Micro Devices • Austin (TX)

On-site
USD 95,000 - 130,000
Health insurance
Paid time off
Professional development
Failure Analysis Engineer
Failure Analysis Engineer

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 160,000
AMD Benefits
Failure Analysis Engineer
Failure Analysis Engineer

AMD • Secaucus (NJ)

On-site
USD 120,000 - 160,000
Benefits at a glance
IT Infrastructure Engineer – RMA & Hardware Diagnostics
IT Infrastructure Engineer – RMA & Hardware Diagnostics

Nebius B.V. • Kansas City (MO)

On-site
USD 90,000 - 130,000
RMA Systems Engineer
RMA Systems Engineer

Vultr • United States

On-site
USD 75,000 - 90,000
Failure Analysis Engineer
Failure Analysis Engineer

FII • San Jose (CA)

On-site
USD 80,000 - 120,000
Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug
Failure Analysis Engineering Manager, GPU ASIC and PCBA Debug

Advanced Micro Devices, Inc. • Secaucus (NJ)

On-site
USD 130,000 - 160,000
Competitive benefits