About the role
We are seeking an experienced RMA Failure Analysis for GPU Servers and enterprise server platforms. The engineer will be responsible for diagnosing, troubleshooting, and performing root cause analysis on customer-returned GPU servers, server motherboards, GPU baseboards, and associated hardware subsystems.
This ideal candidate will possess strong server architecture knowledge, component-level debugging expertise, and the ability to safely handle and analyze high-value hardware throughout the failure analysis process.
What you'll do
- Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, GPU baseboards, and related hardware assemblies.
- Conduct system-level, board-level, and component-level troubleshooting to identify root causes of hardware failures.
- Execute functional testing, diagnostics, and debug activities using standard lab equipment and server validation tools.
- Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
- Analyze failures involving server subsystems including CPUs, GPUs, DIMMs, NICs, SSDs, power supplies, PCIe devices, and cooling/thermal subsystems.
- Troubleshoot hardware issues related to BIOS, BMC, CPLD, FPGA, PCIe, memory, storage, networking, and power delivery circuits.
- Perform component-level debugging including capacitors, resistors, fuses, diodes, MOSFETs, voltage regulators, ICs, and other electronic components.
- Conduct component swapping, isolation testing, and fault reproduction to validate failure mechanisms and root causes.
- Perform detailed visual and mechanical inspections to identify damaged, missing, misaligned, overheated, or improperly assembled components.
- Utilize JIRA and Zendesk to track RMA cases, document failure analysis results, manage issue resolution activities, and maintain clear communication across engineering, quality, and customer support teams.
- Document failure analysis findings, corrective actions, and recommendations to support continuous product quality improvements.
- Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution and corrective actions.
- Follow proper ESD and hardware handling procedures while working with customer-returned products, engineering samples, and production hardware.
Qualifications
Required Qualifications
- 4 years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
- Strong understanding of GPU server architecture and enterprise server platforms.
- Experience performing system-level, board-level, and component-level troubleshooting.
- Ability to read and interpret electrical schematics, block diagrams, and PCB layouts.
- Hands-on experience with server technologies including BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking interfaces.
- Experience using laboratory equipment such as oscilloscopes, digital multimeters (DMM), power analyzers, logic analyzers, and protocol analyzers.
- Working knowledge of Linux operating systems and command-line troubleshooting.
- Strong understanding of root cause analysis methodologies and failure isolation techniques.
- Ability to safely handle sensitive server and GPU hardware while adhering to ESD and hardware handling best practices.
Preferred Qualifications
- Experience supporting AI, HPC, or GPU-accelerated server platforms.
- Experience with customer-returned hardware (RMA) failure analysis processes.
- Knowledge of power delivery architecture, thermal analysis, and signal integrity concepts.
- Familiarity with manufacturing defects, field failures, and reliability-related investigations.
Critical Requirements
- Must be capable of independently troubleshooting GPU servers and server hardware down to the component level.
- Must understand overall server architecture and subsystem interactions before initiating debug activities.
- Must demonstrate strong analytical and problem-solving skills in hardware failure analysis.
- Must be comfortable working with customer-returned hardware and managing multiple RMA investigations simultaneously.
- Must maintain proper hardware handling practices to prevent damage to customer-returned units and engineering samples.
About CoreFleet Solutions
At CoreFleet Solutions, we're building a company focused on delivering exceptional workforce, logistics, and technology services. As a growing startup, every team member has the opportunity to make a meaningful impact and help shape the future of the business.
We partner with organizations to provide staffing solutions, logistics support, and technology deployment services with a commitment to quality, reliability, and customer success.
Why Join CoreFleet
- Opportunity to grow with a fast-growing startup
- Work directly with company leadership
- Learn new skills across multiple industries
- Collaborative, supportive, and entrepreneurial culture
- Make a real impact, your ideas and contributions matter
If you're looking for a place where you can grow your career while helping build something from the ground up, we'd love to hear from you.