Head of Platform Product Reliability

Delos

San Jose (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Housing subsidy
Relocation support
Daily meals
Wellness benefits
Compute budget

Job summary

Etched is seeking a highly technical Head of Platform Product Reliability in San Jose. You will lead reliability engineering across server, rack and datacenter platform products from architecture to fleet deployment.

The role defines reliability strategy, qualification methodologies and accelerated testing programs, partnering with Platform Engineering, Mechanical, Firmware, Manufacturing, Supply Chain and Datacenter Operations to ensure reliability at scale.

Qualifications

  • BS/MS/PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering or related field.
  • 10+ years of reliability engineering experience in hardware-centric organisations.
  • Experience leading reliability programmes for AI accelerators, hyperscale servers or similar platforms.
  • Deep understanding of system‑level failure mechanisms (thermal, power, mechanical).
  • Hands-on with FMEA, Weibull, HALT/HASS and reliability statistics.

Responsibilities

  • Define end-to-end reliability strategy for AI servers, accelerator platforms, rack systems and datacenter infrastructure.
  • Establish reliability requirements, qualification standards and validation methodologies for multiple generations.
  • Build and institutionalise reliability processes across the product life cycle (ALT, AST, environmental testing, power/thermal cycling).
  • Lead root-cause investigations and drive corrective actions across hardware, firmware, thermal and mechanical domains.
  • Develop system reliability models including MTBF, FIT rate analysis and reliability growth tracking.
  • Build fleet reliability infrastructure: telemetry pipelines and field feedback loops to monitor deployed systems.
  • Grow and lead a high-performing reliability engineering team.

Skills

Reliability engineering
System‑level design
Cross‑functional leadership
FMEA
Weibull analysis
HALT/HASS

Education

BS/MS/PhD in Electrical/Mechanical or Reliability Engineering

Job description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software and manufacturing to deliver best‑in‑class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top‑tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

We are seeking a highly technical and execution‑focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack and datacenter platform products.

This role owns system‑level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes and long‑term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems – ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand.

You will work cross‑functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations and Program teams to ensure Etched products achieve exceptional reliability at scale.

Key Responsibilities
  • Define and own the end‑to‑end reliability strategy for AI servers, accelerator platforms, rack systems and datacenter infrastructure, from design requirements through field deployment
  • Establish reliability requirements, qualification standards and validation methodologies that scale across product generations
  • Build and institutionalise reliability engineering processes spanning the full product life‑cycle:
    • EVT / DVT / PVT qualification gates and exit criteria
    • Accelerated life testing (ALT) and accelerated stress testing (AST)
    • Environmental testing: temperature, humidity, altitude, contamination
    • HALT / HASS programmes for design margin and production screening
    • Vibration, shock and transportation stress testing
    • Power cycling, thermal cycling and long‑duration soak testing
  • Lead root‑cause investigations for reliability failures surfaced during development, manufacturing and field deployment, driving corrective actions across hardware, firmware, thermal and mechanical domains
  • Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modelling, component derating methodologies and reliability growth tracking
  • Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change
  • Work closely with ODMs, JDMs, contract manufacturers and component suppliers to validate and enforce long‑term platform reliability commitments
  • Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops and monitoring frameworks that give Etched visibility into deployed system health at scale
  • Drive reliability sign‑off criteria and lead product release readiness reviews across engineering and programme teams
  • Build and lead a high‑performing product reliability engineering organisation – hiring, developing and retaining technical talent as the company scales
You may be a good fit if you have (Must‑have qualifications)
  • BS, MS or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering or a related technical field
  • 10+ years of reliability engineering experience in hardware‑centric organisations, with meaningful time spent on complex systems rather than component‑level work
  • Experience leading reliability programmes for one or more of:
    • AI accelerator or GPU‑class compute systems
    • Hyperscale or cloud server infrastructure
    • Networking platforms, storage systems or rack‑scale infrastructure
  • Deep understanding of system‑level failure mechanisms – including thermal, power delivery, mechanical and connector/interconnect failure modes – and how design decisions affect long‑term field reliability
  • Hands‑on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies and reliability statistics and modelling
  • A track record of driving cross‑functional root‑cause investigations in fast‑moving hardware organisations where schedule pressure is real and accountability is high
  • Strong technical judgment – capable of making defensible trade‑offs between reliability targets, cost, schedule and performance without losing sight of customer expectations
  • Excellent communication skills and the credibility to influence design decisions with engineering leads, programme managers and executive stakeholders
Strong candidates may also have experience with (Nice‑to‑have qualifications)
  • Experience with liquid‑cooled systems, high‑density power delivery or thermal management for high‑power AI infrastructure
  • Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer‑facing reliability commitments and SLA management
  • Demonstrated experience building a reliability organisation from early‑stage – establishing processes, tooling and team norms in environments without established infrastructure
  • Familiarity with fleet telemetry systems, large‑scale field reliability analytics and data‑driven approaches to proactive reliability management
  • Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on‑site qualification engagement
  • Background in high‑speed digital systems, GPU compute platforms or accelerator‑based architectures – with an understanding of how these affect system‑level reliability behaviour
Benefits
  • Medical, dental and vision packages with generous premium coverage
    • $500 per month credit for waiving medical benefits
  • Housing subsidy of $2k per month for those living within walking distance of the office
  • Relocation support for those moving to San Jose (Santana Row)
  • Various wellness benefits covering fitness, mental health and more
  • Daily lunch and dinner in our office
  • Unlimited compute budget subject to ROI justification
How we’re different

Etched believes in the Bitter Lesson. We are the first inference‑focused frontier AI system, betting early on transformer and transformer‑like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in‑person team in San Jose (Santana Row) and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all our technical staff to contribute to both and work across disciplines as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Platform Product Reliability
Head of Platform Product Reliability

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3
Head of Platform Product Reliability
Head of Platform Product Reliability

The Consensus • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+3
Data Center Engineer
Data Center Engineer

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Medical, dental, and vision packages
$500 per month credit for waiving medical benefits
Housing subsidy of $2k per month
+3
Product Engineer, Silicon Validation
Product Engineer, Silicon Validation

The Consensus • San Jose (CA)

On-site
USD 140,000 - 210,000
Medical, dental, and vision packages
Waiver credit of $500 per month for wa
Housing subsidy of $2k per month
+4
Infrastructure Software Engineer
Infrastructure Software Engineer

The Consensus • San Jose (CA)

On-site
USD 190,000 - 240,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+3
Product Quality Engineer
Product Quality Engineer

The Consensus • San Jose (CA)

On-site
USD 140,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support to San Jose
+3
Systems Validation Engineer, L10
Systems Validation Engineer, L10

The Consensus • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for San Jose
+3
Mechanical Operations Lab Lead
Mechanical Operations Lab Lead

The Consensus • San Jose (CA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for San Jose
+3
IT Engineer
IT Engineer

The Consensus • San Jose (CA)

On-site
USD 90,000 - 130,000
Medical, dental, and vision coverage
Housing subsidy
Relocation support (San Jose)
+3
Data Center Engineer
Data Center Engineer

The Consensus • San Jose (CA)

On-site
USD 140,000 - 190,000
Medical, dental, and vision packages
Housing subsidy
Relocation support to San Jose
+3