Chief Architect, Cluster-Scale AI Inference

The Consensus

San Jose (CA)

On-site

USD 260,000 - 520,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision packages
Housing subsidy near Santana Row
Relocation support to San Jose
Wellness benefits
Daily lunch and dinner in office
Unlimited compute budget

Job summary

Etched, building at-scale AI inference supercomputers powered by its own chips, is seeking a Head of Supercomputing to define and lead the architecture, software stack, and operational model for cluster-scale AI compute systems.

This leader will own the end-to-end system software and control-plane strategy—from orchestration, telemetry, provisioning, networking, to fleet reliability—working with ASIC, hardware, kernel, runtime, and infrastructure teams to deliver the highest-performance

Qualifications

  • 15+ years in system software or large-scale compute design.
  • 5+ years leading engineering teams.
  • Deep understanding of hardware/software interfaces (PCIe, RDMA).
  • Experience building or operating cluster-scale AI infrastructure.
  • Proven ability to deliver complex systems from bring-up to production.
  • Strong leadership and cross-functional collaboration.

Responsibilities

  • Define and drive the technical vision and roadmap for Etched’s Supercomputing software stack, from node bring-up to multi-rack clusters.
  • Build, scale, and lead a high-performance Supercomputing organization.
  • Directly manage and develop 15+ engineers with a variety of experience levels.
  • Architect and own low-level control-plane software for system bring-up, provisioning, networking, configuration, and fleet management.
  • Define orchestration primitives for managing devices, nodes, racks, and full cluster deployments.
  • Oversee development of system services that interface with firmware, drivers, kernel subsystems, and runtime layers.
  • Establish system telemetry and observability infrastructure for customers.
  • Own the system software lifecycle from first silicon bring-up through stable production releases.
  • Collaborate with manufacturing and test engineering to integrate diagnostics and system software into factory environments.
  • Define reliability targets, operational metrics, and release processes for production deployments.
  • Recruit, mentor, and retain exceptional systems engineers and engineering leaders.
  • Act as a senior technical voice shaping company-wide infrastructure decisions.

Skills

System software leadership
Large-scale compute systems
Hardware/software interfaces
Team leadership
HPC/AI infrastructure
PCIe/RDMA knowledge
Debugging HW-SW
Cross-functional collaboration

Job description

Etched, building at-scale AI inference supercomputers powered by its own chips, is seeking a Head of Supercomputing to define and lead the architecture, software stack, and operational model for cluster-scale AI compute systems.

This leader will own the end-to-end system software and control-plane strategy—from orchestration, telemetry, provisioning, networking, to fleet reliability—working with ASIC, hardware, kernel, runtime, and infrastructure teams to deliver the highest-performance

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Chief Architect, AI Inference Supercomputing
Chief Architect, AI Inference Supercomputing

Etched • San Jose (CA)

On-site
USD 180,000 - 230,000
Medical, dental, and vision packages
$500 per month credit for waiving medical benefits
Housing subsidy of $2k per month
+2
Head of AI Inference Supercomputing
Head of AI Inference Supercomputing

Etched.ai, Inc. • San Jose (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+1
Cluster-Scale AI Systems Engineer
Cluster-Scale AI Systems Engineer

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical coverage
Dental coverage
Vision coverage
+4
Lead Architect, Scaled AI Inference & Systems
Lead Architect, Scaled AI Inference & Systems

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Senior Architect, Scaled AI Inference & Orchestration
Senior Architect, Scaled AI Inference & Orchestration

NVIDIA • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Head of Supercomputing
Head of Supercomputing

The Consensus • San Jose (CA)

On-site
USD 260,000 - 520,000
Medical, dental, and vision packages
Housing subsidy near Santana Row
Relocation support to San Jose
+3
Senior Architect, Scaled AI Inference & Systems — Equity
Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Head of Supercomputing
Head of Supercomputing

Etched.ai, Inc. • San Jose (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+1