Hardware Diagnostics Engineer: Burn-In, RMAs & Debugging

Tensorwave

Las Vegas (NV)

On-site

USD 110,000 - 170,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock Options
Medical, Dental, Vision insurance 100%
Health Savings Account contributions
Disability Insurance
Life Insurance
Paid Holidays
Flexible PTO
Parental Leave
401(k)
Employee Assistance Program
On-site amenities

Job summary

TensorWave is seeking a Hardware Diagnostics Engineer to run burn-in tests, triage failures, and manage RMAs end-to-end. You will work with IPMI/Redfish for out-of-band control and ensure fleet reliability across GPU servers.

If you enjoy diagnosing hardware issues and documenting evidence, this role suits you. Ideal candidates have 3–6 years in datacenter operations and hands-on server hardware experience, including BMC management and Linux troubleshooting.

Qualifications

  • 3-6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering.
  • Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, and reading hardware behavior at the rack.
  • Practical experience with BMCs and out-of-band management: IPMI, Redfish, iDRAC, iLO, or equivalent.
  • Strong Linux troubleshooting: boot process, driver and device issues, and diagnostic tools such as dmesg, lspci, ipmitool, and SMART.
  • Comfort reading sensor data, event logs, and telemetry to identify failures.
  • Working scripting ability in Bash or Python to automate tasks.
  • Experience running hardware RMAs with vendors or driving issues to closure with outside parties.
  • Methodical troubleshooting: isolate variables and document evidence.

Responsibilities

  • Run server and GPU burn-in and stress testing; interpret results to decide production readiness.
  • Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling; reproduce, isolate, and document fault.
  • Work servers out-of-band via IPMI and Redfish for power, BIOS, and sensor data collection.
  • Apply firmware updates across fleet following baselines and rollout process.
  • Drive RMAs with vendors from ticket to replacement and return of failed parts.
  • Keep NetBox asset, serial, and replacement history accurate for every rack.
  • Track failure patterns across fleet and flag recurring parts/firmware versions.
  • Improve runbooks and automate repetitive steps; partner with data-center operations during turn-ups.
  • Participate in on-call and escalation rotation for hardware issues.

Skills

Datacenter operations
Linux troubleshooting
Scripting (bash/python)
Vendor RMAs management
Written communication
System diagnostics

Tools

IPMI
Redfish
iDRAC
iLO
NetBox
Ansible
Python (REST APIs)
dmesg

Job description

TensorWave is seeking a Hardware Diagnostics Engineer to run burn-in tests, triage failures, and manage RMAs end-to-end. You will work with IPMI/Redfish for out-of-band control and ensure fleet reliability across GPU servers.

If you enjoy diagnosing hardware issues and documenting evidence, this role suits you. Ideal candidates have 3–6 years in datacenter operations and hands-on server hardware experience, including BMC management and Linux troubleshooting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Hardware Diagnostics Engineer - Infrastructure
Hardware Diagnostics Engineer - Infrastructure

Tensorwave • Las Vegas (NV)

On-site
USD 110,000 - 170,000
Stock Options
Medical, Dental, Vision insurance 100%
Health Savings Account contributions
+8
RMA Failure Analysis Engineer
RMA Failure Analysis Engineer

Core Fleet Solutions • San Jose (CA)

On-site
USD 120,000 - 170,000
Hardware Diagnostics Engineer — Remote, Equity
Hardware Diagnostics Engineer — Remote, Equity

MatX Inc. • Mountain View (CA)

Hybrid
USD 160,000 - 600,000
4 weeks PTO (accrued)
Remote work up to 3 weeks
Health insurance
+1
Senior GPU Server Failure Analyst (RMA & Root Cause)
Senior GPU Server Failure Analyst (RMA & Root Cause)

CoreFleet Solutions • San Jose (CA)

On-site
USD 90,000 - 150,000
Senior Diagnostics Platform Engineer - Data Center HW/SW
Senior Diagnostics Platform Engineer - Data Center HW/SW

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Data Center Hardware Test Technician – Bring-Up Specialist
Data Center Hardware Test Technician – Bring-Up Specialist

Programmers.io • Houston (TX)

On-site
USD 55,000 - 75,000
Hardware Reliability Engineer – HTOL & Burn-In
Hardware Reliability Engineer – HTOL & Burn-In

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD Benefits
Firmware-Savvy Failure Analysis Engineer - Server Platforms
Firmware-Savvy Failure Analysis Engineer - Server Platforms

AMD • Secaucus (NJ)

On-site
USD 130,000 - 190,000
Benefits at a glance
Server Test Technician - BIOS/OS & Hardware Diagnostics
Server Test Technician - BIOS/OS & Hardware Diagnostics

Raso360 • Fremont (CA)

On-site
USD 52,000 - 75,000
Senior Data Center GPU Validation & Debug Engineer
Senior Data Center GPU Validation & Debug Engineer

AMD • Austin (TX)

Hybrid
USD 120,000 - 170,000
AMD Benefits at a glance