Performance Tools Intern

The Consensus

San Jose (CA)

On-site

USD 34,000 - 62,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Etched in San Jose is seeking a talented intern to design and develop a performance analysis tool for our ML accelerator. You will build tooling to understand workload behavior, identify bottlenecks, and help optimize hardware and software interactions.

Work with hardware, compiler, firmware, and inference engineers in an in-person setting. Strong coding skills in C++ or Rust, plus Python, will help you contribute from day one.

Qualifications

  • Proficient in C++ or Rust with strong systems programming background.
  • Interest or experience in low-level performance analysis and profiling.
  • Familiarity with hardware performance counters, traces, and CPUs/accelerators.

Responsibilities

  • Build components of our performance analysis and profiling infrastructure.
  • Collect and analyze performance data from custom ML accelerators, including hardware counters, execution traces, and memory behavior.
  • Develop tooling to trace host-side runtime activity, system behavior, and accelerator execution.
  • Help correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads.
  • Build analysis and visualization tools that help engineers identify performance bottlenecks and optimize models.
  • Work alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and develop tools that improve developer productivity.

Skills

C++
Rust
Python

Tools

Nsight
VTune
Perfetto
Xprof

Job description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

Join our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.

During your internship, you may:
  • Build components of our performance analysis and profiling infrastructure.
  • Collect and analyze performance data from our custom ML accelerators, including hardware counters, execution traces, and memory behavior.
  • Develop tooling to trace host-side runtime activity, system behavior, and accelerator execution.
  • Help correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads.
  • Build analysis and visualization tools that help engineers identify performance bottlenecks and optimize models.
  • Work alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and develop tools that improve developer productivity.
Representative projects
  • Implement the data collection framework for hardware performance counters on a custom PCIe-based accelerator.
  • Develop a user-space service for low-overhead tracing of accelerator activity.
  • Design and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.
  • Create an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.
You may be a good fit if you have
  • Strong programming skills in C++ or Rust. Experience with Python is a plus.
  • Solid understanding of computer architecture, including CPUs, GPUs or AI accelerators, memory hierarchies, and parallel programming.
  • Experience or strong interest in low-level performance analysis, profiling, and performance optimization.
  • Familiarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto, or similar tools is a plus.
  • Experience or strong interest in operating systems, compilers, firmware, drivers, or other low-level systems software.
  • Passion for understanding how complex systems behave under real workloads and building tools that help other engineers optimize performance.
  • Strong problem-solving skills and curiosity to learn quickly in a fast-paced engineering environment.
Strong candidates may also have experience with (Nice-to-have qualifications)
  • Direct experience developing performance analysis or debugging tools.
  • Experience with ML accelerator architectures (GPUs, TPUs, etc.).
  • Experience with kernel-mode driver development (Linux or Windows).
How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Tools Intern
Performance Tools Intern

Etched • San Jose (CA)

On-site
USD 20,000 - 31,000
Head of Performance Visibility
Head of Performance Visibility

The Consensus • San Jose (CA)

On-site
USD 180,000 - 260,000
Medical/dental/vision
Housing subsidy
Relocation assistance
+3
Software Engineer – Performance Profiling
Software Engineer – Performance Profiling

The Consensus • San Jose (CA)

On-site
USD 180,000 - 260,000
Medical coverage
Housing subsidy
Relocation support
+3
Performance Characterization Engineer
Performance Characterization Engineer

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+1
Software Engineer – Performance Profiling
Software Engineer – Performance Profiling

Etched • San Jose (CA)

On-site
USD 150,000 - 275,000
Medical, dental, and vision packages
$500 monthly credit for waiving medical benefits
$2k housing subsidy
+2
Software Engineer – Performance Profiling
Software Engineer – Performance Profiling

Delos • San Jose (CA)

On-site
USD 150,000 - 275,000
Full medical, dental, and vision packages
Housing subsidy of $2,000/month
Daily lunch and dinner in office
+1
Head of Performance Visibility
Head of Performance Visibility

Etched • San Jose (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision packages with generous coverage
Housing subsidy of $2k per month
Relocation support
+1
Performance Modeling Engineer
Performance Modeling Engineer

The Consensus • San Jose (CA)

On-site
USD 150,000 - 230,000
Medical, dental, and vision coverage
Housing subsidy near Santana Row: $2k/
Relocation support to San Jose
+3
Inference Intern
Inference Intern

Etched • San Jose (CA), Northern (KY)

Hybrid
USD 13,225,000 - 19,837,000
Housing support
Lunch & dinner
Mentorship
+1
Head of Performance Visibility
Head of Performance Visibility

Delos • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for new hires
+2