AI Hardware Reliability Engineer: Pod-Level RAS Telemetry

Intel

Santa Clara (CA)

On-site

USD 122,000 - 232,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock bonuses
Health benefits
On-site amenities

Job summary

Intel is seeking an experienced reliability engineer to define and own pod-level reliability specs across compute, memory, storage, network, power and cooling for AI hardware data centers.

You will lead FMEA, telemetry, and KPI tracking; collaborate with facilities on N+1 redundancy and disaster recovery, and apply Weibull/FIT reliability methods. AI cluster operations and data analytics are a plus.

Qualifications

  • BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.
  • Experience authoring and owning reliability specs and requirement flow-down.
  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.
  • Experience with large-scale fleet telemetry and thermal/power redundancy.

Responsibilities

  • Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.
  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.
  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.
  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.
  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.
  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Skills

AI cluster operations
Data analytics
Python
SQL

Education

BS/MS/PhD in EE/ME Reliability or related

Job description

Intel is seeking an experienced reliability engineer to define and own pod-level reliability specs across compute, memory, storage, network, power and cooling for AI hardware data centers.

You will lead FMEA, telemetry, and KPI tracking; collaborate with facilities on N+1 redundancy and disaster recovery, and apply Weibull/FIT reliability methods. AI cluster operations and data analytics are a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer - AI Hardware & Data Center RAS
Reliability Engineer - AI Hardware & Data Center RAS

Intel • Boxborough (MA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
Retirement plans
+1
Reliability Engineer
Reliability Engineer

Intel • Boxborough (MA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
Retirement plans
+1
Reliability Engineer
Reliability Engineer

Intel • Santa Clara (CA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
On-site amenities
Lead AI/HPC Cluster Reliability Architect
Lead AI/HPC Cluster Reliability Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 240,000
AMD benefits at a glance
Hardware Analytics Engineer: AI Telemetry & Reliability
Hardware Analytics Engineer: AI Telemetry & Reliability

Cerebras • United States

Remote
USD 120,000 - 190,000
Rack-Level Reliability Engineer for HPC & AI Systems
Rack-Level Reliability Engineer for HPC & AI Systems

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 120,000 - 160,000
Chief Platform Reliability Architect for AI Infrastructure
Chief Platform Reliability Architect for AI Infrastructure

The Consensus • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+3
Senior Reliability Engineer, AI Server Hardware
Senior Reliability Engineer, AI Server Hardware

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 140,000 - 170,000
Hardware Reliability Engineer – AI Accelerator Packaging
Hardware Reliability Engineer – AI Accelerator Packaging

DensityAI • Mountain View (CA)

On-site
USD 220,000 - 350,000
Lead Architect, AI/HPC Cluster Reliability
Lead Architect, AI/HPC Cluster Reliability

AMD • Austin (TX)

On-site
USD 180,000 - 260,000