Head of AI Accelerator Performance & Telemetry

The Consensus

San Jose (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical/dental/vision
Housing subsidy
Relocation assistance
Wellness programs
Daily meals
Compute budget

Job summary

Etched is hiring a Head of Performance Visibility to define performance metrics and abstractions for next-generation AI accelerator systems. This role spans hardware and software, designing cross-layer telemetry, counters, and distributed reasoning across nodes, racks, and clusters.

You will shape scalable tooling and insights for engineers debugging large-scale AI systems. We operate in a fully in-person team in San Jose (Santana Row) and seek someone who can drive architecture across hardware

Qualifications

  • Experience building profiling or observability systems.
  • Ability to translate raw hardware signals into scalable telemetry and analysis infrastructure.
  • Experience correlating time-series events across distributed systems.
  • Deep systems programming experience (C++ or Rust).
  • Experience designing distributed tracing or performance modeling frameworks.

Responsibilities

  • Define the architectural approach for collecting and structuring telemetry across CPUs, drivers, interconnects, and multiple accelerators
  • Design scalable models for correlating performance events across device and host boundaries
  • Develop mechanisms to align hardware counters, runtime activity, and workload semantics across model-layer execution into coherent insight
  • Implement time synchronization and trace-alignment strategies across multi-device systems
  • Build tools that identify bottlenecks among multi-accelerator workloads across chips within hosts
  • Contribute to analysis engines and developer-facing tooling that transform raw telemetry into intuitive insight

Skills

Systems programming
Distributed systems
C++/Rust
Telemetry/observability
Time-series analysis

Job description

Etched is hiring a Head of Performance Visibility to define performance metrics and abstractions for next-generation AI accelerator systems. This role spans hardware and software, designing cross-layer telemetry, counters, and distributed reasoning across nodes, racks, and clusters.

You will shape scalable tooling and insights for engineers debugging large-scale AI systems. We operate in a fully in-person team in San Jose (Santana Row) and seek someone who can drive architecture across hardware

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Performance Visibility
Head of Performance Visibility

The Consensus • San Jose (CA)

On-site
USD 180,000 - 260,000
Medical/dental/vision
Housing subsidy
Relocation assistance
+3
AI Accelerator Performance Profiling Architect
AI Accelerator Performance Profiling Architect

Etched • San Jose (CA)

On-site
USD 200,000 - 300,000
Head of AI Performance Profiling & Visibility
Head of AI Performance Profiling & Visibility

Delos • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for new hires
+2
Head of Performance Visibility
Head of Performance Visibility

Etched • San Jose (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision packages with generous coverage
Housing subsidy of $2k per month
Relocation support
+1
Head of Performance Visibility
Head of Performance Visibility

Delos • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for new hires
+2
Hardware Systems Engineer, AI Accelerators (100G PCIe)
Hardware Systems Engineer, AI Accelerators (100G PCIe)

The Consensus • San Jose (CA)

On-site
USD 150,000 - 275,000
Medical, dental, and vision coverage
Housing subsidy near office
Relocation support
+3
Performance Analysis Lead for ML Accelerator Hardware
Performance Analysis Lead for ML Accelerator Hardware

The Consensus • San Jose (CA)

On-site
USD 180,000 - 260,000
Medical coverage
Housing subsidy
Relocation support
+3
Senior AI Cluster Performance Architect & Telemetry Lead
Senior AI Cluster Performance Architect & Telemetry Lead

NVIDIA • Austin (TX)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Cluster Performance & Telemetry Architect
Senior AI Cluster Performance & Telemetry Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits package
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model