Reliability Engineer - AI Hardware & Data Center RAS

Intel

Boxborough (MA)

On-site

USD 122,000 - 232,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock bonuses
Health benefits
Retirement plans
Vacation

Job summary

Intel is seeking an experienced Reliability Engineer to define and own pod-level reliability specifications for a large-scale AI hardware data center. You will drive RAS-focused design targets across compute, memory, storage, network, and power subsystems.

You will lead FMEA, root-cause analysis, and telemetry-driven improvements, partnering with facilities on N+1 redundancy and disaster recovery. This on-site role at multiple US locations emphasizes data-driven reliability across silicon and

Qualifications

  • Minimum: BS/MS/PhD in EE/ME Reliability or related; 4–6 yrs experience.
  • Experience authoring and owning reliability specs and requirement flow-down.
  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.
  • Experience with large-scale fleet telemetry and thermal/power redundancy.

Responsibilities

  • Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.
  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.
  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.
  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.
  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.
  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Skills

RAS
FMEA
Statistical reliability
Fleet telemetry
Thermal/power redundancy
Python
SQL

Education

BS/MS/PhD in EE/ME Reliability or related

Job description

Intel is seeking an experienced Reliability Engineer to define and own pod-level reliability specifications for a large-scale AI hardware data center. You will drive RAS-focused design targets across compute, memory, storage, network, and power subsystems.

You will lead FMEA, root-cause analysis, and telemetry-driven improvements, partnering with facilities on N+1 redundancy and disaster recovery. This on-site role at multiple US locations emphasizes data-driven reliability across silicon and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Hardware Reliability Engineer: Pod-Level RAS Telemetry
AI Hardware Reliability Engineer: Pod-Level RAS Telemetry

Intel • Santa Clara (CA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
On-site amenities
Reliability Engineer
Reliability Engineer

Intel • Boxborough (MA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
Retirement plans
+1
Reliability Engineer
Reliability Engineer

Intel • Santa Clara (CA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
On-site amenities
Chief Platform Reliability Architect for AI Infrastructure
Chief Platform Reliability Architect for AI Infrastructure

The Consensus • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+3
Hardware Reliability Engineer – AI Data Center SRE
Hardware Reliability Engineer – AI Data Center SRE

SpaceXAI • Southaven (MS)

On-site
USD 110,000 - 160,000
On-site Memphis role
Platform Reliability Leader for AI Infrastructure
Platform Reliability Leader for AI Infrastructure

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3
Lead AI/HPC Cluster Reliability Architect
Lead AI/HPC Cluster Reliability Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 240,000
AMD benefits at a glance
Quality & Reliability Engineer, AI Power Modules
Quality & Reliability Engineer, AI Power Modules

Renesas Electronics • Illinois

On-site
USD 140,000 - 200,000
Hardware Reliability Engineer – AI Accelerator Packaging
Hardware Reliability Engineer – AI Accelerator Packaging

DensityAI • Mountain View (CA)

On-site
USD 220,000 - 350,000
Senior Reliability Engineer: AI-Powered Systems
Senior Reliability Engineer: AI-Powered Systems

Aivres • Fremont (CA)

Hybrid
USD 150,000 - 250,000
Competitive salary and benefits package
Opportunity to work on cutting-edge technology
Career development and growth opportunities