Turn this role into an interview — a resume and cover letter built around what this employer wants.
Delos is hiring a Head of Performance Profiling to define performance across next-generation AI accelerator systems. This key role will shape how performance intelligence is utilized in complex hardware and software environments.
The candidate must possess deep systems programming expertise, particularly in C++ or Rust, and experience in distributed systems. The position offers various benefits including medical coverage and relocation support for those moving to San Jose.
San Jose
Full time
On-site
Software
Etched is building the world’s first AI inference system purpose-built for transformers - delivering over 10x higher performance and dramatically lower cost and latency than a B200. With Etched ASICs, you can build products that would be impossible with GPUs, like real‑time video generation models and extremely deep & parallel chain‑of‑thought reasoning agents. Backed by hundreds of millions from top‐tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.
We are hiring a Head of Performance Profiling to define how performance is understood across next‑generation AI accelerator systems.
Our ML accelerator platform spans custom silicon, supercomputing software, compiler stacks, runtime libraries, and distributed inference environments. Performance at this scale is no longer a device‑level question — it is a high‑performance distributed system problem. You will define the performance metrics that connect raw hardware signals to distributed workload context, ML cluster dynamics, pod communication patterns, and emergent bottlenecks.
This role requires more than telemetry. You will establish new abstractions, structured counter ontologies, cross‑layer event correlation frameworks, distributed time‑alignment strategies, and scalable reasoning systems operating across nodes, racks, and clusters. Working at the intersection of hardware design, driver architecture, runtime systems, and ML infrastructure, you will shape how these layers expose and consume performance intelligence. This is a foundational role defining not just tooling, but how our platform reasons about efficiency, scalability, and system behavior for years to come.
Etched believes in the Bitter Lesson. We think most of the progress in the AI field has come from using more FLOPs to train and run models, and the best way to get more FLOPs is to build model‑specific hardware. Larger and larger training runs encourage companies to consolidate around fewer model architectures, which creates a market for single‑model ASICs.
We are a fully in‑person team in San Jose and Taipei, and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both as needed.