Chief Platform Reliability Architect for AI Infrastructure

The Consensus

San Jose (CA)

On-site

USD 210,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision packages
Housing subsidy
Relocation support
Wellness benefits
Daily lunch and dinner
Unlimited compute budget

Job summary

Etched is seeking a highly technical Head of Platform Product Reliability to lead reliability engineering across server, rack, and datacenter platforms. You will define reliability strategies, qualification methodologies, and long-term standards for complex AI infrastructure systems.

You will work with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, and Supply Chain to ensure reliability at scale, driving RCAs, and building fleet telemetry and monitoring.

Qualifications

  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field.
  • 10+ years of reliability engineering experience in hardware-centric organizations, with focus on complex systems.

Responsibilities

  • Define end-to-end reliability strategy for AI servers, accelerators, rack systems, and datacenter infrastructure.
  • Establish reliability requirements, qualification standards, and validation methodologies across product generations.
  • Lead reliability programs and conduct RCIs to drive corrective actions across hardware, firmware, and mechanical domains.
  • Develop system reliability models including MTBF, FIT, Weibull, and derating methodologies.
  • Build fleet reliability infrastructure: telemetry, field feedback, and monitoring for deployed systems.
  • Lead a high-performing reliability engineering organization and mentor technical talent.

Skills

Reliability engineering
Cross-functional collaboration
Analytical reasoning
Communication

Education

BS/MS/PhD in Electrical or Mechanical Engineering or Reliability Engineering

Tools

FMEA
Weibull analysis
HALT/HASS
Failure analysis

Job description

Etched is seeking a highly technical Head of Platform Product Reliability to lead reliability engineering across server, rack, and datacenter platforms. You will define reliability strategies, qualification methodologies, and long-term standards for complex AI infrastructure systems.

You will work with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, and Supply Chain to ensure reliability at scale, driving RCAs, and building fleet telemetry and monitoring.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Reliability Leader for AI Infrastructure
Platform Reliability Leader for AI Infrastructure

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3
Senior Platform Reliability Lead for AI Servers, Datacenters
Senior Platform Reliability Lead for AI Servers, Datacenters

Delos • San Jose (CA)

On-site
USD 180,000 - 280,000
Housing subsidy
Relocation support
Daily meals
+2
Head of Platform Product Reliability
Head of Platform Product Reliability

Delos • San Jose (CA)

On-site
USD 180,000 - 280,000
Housing subsidy
Relocation support
Daily meals
+2
Head of Platform Product Reliability
Head of Platform Product Reliability

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3
Staff SRE: Reliability Architect for AI-Driven Platform
Staff SRE: Reliability Architect for AI-Driven Platform

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Benefits package
Reliability Engineer - AI Hardware & Data Center RAS
Reliability Engineer - AI Hardware & Data Center RAS

Intel • Boxborough (MA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
Retirement plans
+1
Platform Engineer: AI-Driven Infra for Scale & Reliability
Platform Engineer: AI-Driven Infra for Scale & Reliability

Alldus International Consulting Ltd • San Francisco (CA)

On-site
USD 200,000 - 250,000
Comprehensive benefits package
AI Reliability Engineering Lead — Production-Grade AI
AI Reliability Engineering Lead — Production-Grade AI

Socket.dev • Cincinnati (OH)

On-site
USD 180,000 - 240,000
Senior Backend Engineer - AI-Driven Reliability Platform
Senior Backend Engineer - AI-Driven Reliability Platform

Affirm • Miami (FL)

On-site
USD 173,000 - 233,000
Health and wellness benefits
Remote-first culture
Competitive equity
Remote Backend Reliability Engineer - AI-Driven Platform
Remote Backend Reliability Engineer - AI-Driven Platform

Affirm • Boulder (CO)

On-site
USD 173,000 - 255,000
Health care coverage
Flexible Spending Wallets
Time off
+1