Data Center L2 Hardware & GPU Debug Engineer

Servicio Industrial de Comedores SA de CV

Santiago de Querétaro

Presencial

MXN 450.000 - 750.000

Jornada completa

hace 17 horas
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Servicio Industrial de Comedores SA de CV in Mexico seeks an experienced L2 Hardware & Failure Analysis Engineer to support mission-critical enterprise server hardware, including x86 servers and AI/GPU platforms.

Responsibilities include hands-on Break/Fix, failure analysis, and log extraction via BMC/IPMI/Redfish. You will diagnose PCIe, iSCSI, RoCE, SAS, Fibre Channel issues, perform post-repair validations, and author RCA reports. Fluent English is mandatory for global escalations.

Formación

  • 5+ years of hands-on experience in enterprise server hardware troubleshooting, validation, or FA.
  • Fluent English required for global escalations.
  • Deep knowledge of x86 architecture, Linux CLI, and server management.
  • Direct experience with GPU servers (NVIDIA / AMD) is a big plus.

Responsabilidades

  • Perform Break/Fix and L10/L11 failure analysis on enterprise x86 servers and AI/GPU platforms.
  • Extract and analyze low-level event logs using BMC, IPMI, and Redfish interfaces.
  • Diagnose PCIe, iSCSI, RoCE, SAS, Fibre Channel issues in servers and storage controllers.
  • Execute post-repair validations, firmware flashing (BIOS/BMC), and quality checks.
  • Write Root Cause Analysis (RCA) reports, SOPs, and KB articles.

Conocimientos

Break/Fix
x86 Architecture
Linux CLI
GPU Servers
BMC/IPMI/Redfish
RCA & SOP documentation
Enterprise server hardware
Troubleshooting

Descripción del empleo

Are you passionate about enterprise server hardware, x86 architecture, and cutting-edge AI / GPU platforms? We are looking for an experienced L2 Hardware & Failure Analysis Engineer to join our team supporting mission-critical environments!

Language: Fluent English (Mandatory for global escalations)

  • Perform hands-on Break/Fix and L10/L11 failure analysis on enterprise x86 servers and AI/GPU platforms (CPU, DDR5, Motherboards, PSUs, PCIe).
  • Extract and analyze low-level event logs using BMC, IPMI, and Redfish interfaces.
  • Diagnose network & storage controller issues (PCIe, iSCSI, RoCE, SAS, Fibre Channel).
  • Execute post-repair validations, firmware flashing (BIOS/BMC), and quality checks in production/lab environments.
  • Write Root Cause Analysis (RCA) reports, SOPs, and KB articles.
  • 5+ years of hands-on experience in enterprise server hardware troubleshooting, validation, or Failure Analysis (FA).
  • Deep knowledge of x86 architecture, Linux OS command line, and server management.
  • Direct experience with GPU Servers (NVIDIA / AMD platforms) is a massive plus!
  • Experience in companies like Intel, Foxconn, Ingrasys, Jabil, Wistron, Oracle, Microsoft, or hyperscale Data Centers.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Data Center L2 Hardware & GPU Failure Analyst
Data Center L2 Hardware & GPU Failure Analyst

Servicio Industrial de Comedores SA de CV • Santiago de Querétaro

Presencial
MXN 450.000 - 750.000
Lead Test Development Engineer
Lead Test Development Engineer

Jabil • Región Centro

Híbrido
MXN 600.000 - 900.000
Linux Systems Engineer
Linux Systems Engineer

Luxoft • Región Centro

Presencial
MXN 300.000 - 450.000
Test Engineer
Test Engineer

Lenovo • Apodaca

Presencial
MXN 400.000 - 560.000
Senior Storage Systems Engineer
Senior Storage Systems Engineer

Luxoft • Región Centro

Presencial
MXN 600.000 - 900.000
Failure Analysis Technician
Failure Analysis Technician

Jabil • Región Centro

Presencial
MXN 433.000 - 608.000
HSIO/DDR Bench Test Engineer
HSIO/DDR Bench Test Engineer

North American Production Sharing, Inc. • Tijuana

Presencial
MXN 600.000 - 900.000
Data Center Operations Engineer III — GPU Infra
Data Center Operations Engineer III — GPU Infra

Lambda • Santiago de Querétaro

Presencial
MXN 531.000 - 708.000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Failure Analysis Engineer III
Failure Analysis Engineer III

Jabil • Chihuahua

Presencial
MXN 480.000 - 640.000
Test Engineer
Test Engineer

Lenovo • Monterrey

Presencial
MXN 699.000 - 1.050.000