Head of Inference Performance & System Visibility

Etched

San Jose (CA)

On-site

USD 240,000 - 340,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision
Housing subsidy
Relocation support
Wellness benefits
Daily lunch and dinner
Unlimited compute budget

Job summary

Etched in San Jose is seeking a Head of Performance Visibility to define how performance is understood across next-generation AI accelerator systems. You will establish abstractions, cross-layer event correlation, and scalable reasoning across nodes and racks, shaping tooling and telemetry for years to come.

You will lead engineering across hardware, drivers, runtimes, and ML infrastructure, driving performance intelligence, time-alignment strategies, and production-grade telemetry for

Qualifications

  • Deep experience building complex systems at the intersection of hardware and software.
  • Personally envisioned and built significant portions of profiling, tracing, or observability systems.
  • Demonstrated ability to translate raw hardware signals into scalable telemetry and analysis infrastructure.
  • Experience correlating time-series events across distributed systems.
  • Deep systems programming expertise (C++ or Rust) near hardware or runtime systems.
  • Experience designing distributed correlation mechanisms, timestamp-alignment, or performance modeling frameworks.
  • Experience leading cross-functional architectural initiatives spanning hardware and software teams.

Responsibilities

  • Define the architectural approach for performance telemetry across CPUs, drivers, interconnects, and accelerators.
  • Design scalable models correlating performance events across device and host boundaries.
  • Develop mechanisms to align hardware counters, runtime activity, and workload semantics for coherent insight.
  • Implement time synchronization and trace-alignment across multi-device systems.
  • Define structured counter taxonomies and derived performance models.
  • Build tools to identify bottlenecks across multi-accelerator workloads and distributed inference.

Skills

Hardware-software co-design
Profiling/Observability
C++ or Rust
Distributed tracing
Time-series correlation
Performance modeling
Leadership

Tools

Telemetry tooling
Distributed systems tooling

Job description

Etched in San Jose is seeking a Head of Performance Visibility to define how performance is understood across next-generation AI accelerator systems. You will establish abstractions, cross-layer event correlation, and scalable reasoning across nodes and racks, shaping tooling and telemetry for years to come.

You will lead engineering across hardware, drivers, runtimes, and ML infrastructure, driving performance intelligence, time-alignment strategies, and production-grade telemetry for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of AI Accelerator Performance & Telemetry
Head of AI Accelerator Performance & Telemetry

The Consensus • San Jose (CA)

On-site
USD 180,000 - 260,000
Medical/dental/vision
Housing subsidy
Relocation assistance
+3
ML Inference Performance Visibility Engineer
ML Inference Performance Visibility Engineer

Etched • San Jose (CA)

On-site
USD 150,000 - 210,000
Housing subsidy
Relocation support
Medical benefits
+1
Head of Performance Visibility
Head of Performance Visibility

The Consensus • San Jose (CA)

On-site
USD 180,000 - 260,000
Medical/dental/vision
Housing subsidy
Relocation assistance
+3
Head of AI Performance Profiling & Visibility
Head of AI Performance Profiling & Visibility

Delos • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for new hires
+2
Head of Inference Performance Visibility
Head of Inference Performance Visibility

Etched • San Jose (CA)

On-site
USD 240,000 - 340,000
Medical, dental, and vision
Housing subsidy
Relocation support
+3
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Infrastructure Performance & Modeling Lead
AI Infrastructure Performance & Modeling Lead

Renice AI • Mountain View (CA)

Hybrid
USD 180,000 - 280,000
Senior AI Infra Performance & Observability Engineer
Senior AI Infra Performance & Observability Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical insurance
Life Insurance
401(k) with match
+2
Head of AI Inference Supercomputing
Head of AI Inference Supercomputing

Etched.ai, Inc. • San Jose (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+1
Head of Performance Visibility
Head of Performance Visibility

Delos • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for new hires
+2