Hardware Analytics Engineer: AI Telemetry & Reliability

Cerebras

United States

Remote

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems is seeking a Hardware Analytics Engineer to design and optimize scalable data pipelines for multi-terabyte hardware telemetry and reliability analytics. You will architect ETL processes to aggregate performance streams and build anomaly detection systems using Python, SQL, Tableau, Hive, and Spark.

You will lead hardware characterization experiments and thermal studies to drive improvements in reliability and efficiency across CPU, GPU, DRAM, PCIe, and storage subsystems.

Qualifications

  • Master’s degree or foreign equivalent in Electrical Engineering, Computer Engineering, or closely related field.

Responsibilities

  • Design scalable data pipelines for multi-terabyte telemetry, reliability analytics, and performance optimization.
  • Architect and optimize ETL processes to aggregate and analyze hardware performance and telemetry streams across platforms.
  • Develop hardware performance analysis and anomaly detection using Python, SQL, Tableau, Hive and Spark to forecast failures and optimize systems.
  • Lead hardware characterization experiments and thermal studies to improve reliability and efficiency, reducing carbon footprint and water usage.
  • Engineer telemetry ingestion, monitoring, and visualization to provide real-time health data to operations teams for data-driven decisions.

Skills

Python
SQL
Tableau
Hive
Spark
ETL
Data pipelines
Telemetry
Anomaly detection

Education

Master’s degree in Electrical Engineering

Job description

Cerebras Systems is seeking a Hardware Analytics Engineer to design and optimize scalable data pipelines for multi-terabyte hardware telemetry and reliability analytics. You will architect ETL processes to aggregate performance streams and build anomaly detection systems using Python, SQL, Tableau, Hive, and Spark.

You will lead hardware characterization experiments and thermal studies to drive improvements in reliability and efficiency across CPU, GPU, DRAM, PCIe, and storage subsystems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Hardware Analytics Engineer: AI Server Telemetry
Remote Hardware Analytics Engineer: AI Server Telemetry

Cerebras • Sunnyvale (CA)

On-site
USD 213,000 - 225,000
Telecommuting permitted
Hardware Analytics Engineer
Hardware Analytics Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 213,000 - 225,000
Telecommuting permitted
Engineering Manager, Kernel Reliability
Engineering Manager, Kernel Reliability

Cerebras • United States

On-site
USD 120,000 - 160,000
Engineering Manager, Kernel Reliability
Engineering Manager, Kernel Reliability

Cerebras • Raleigh (NC)

On-site
USD 180,000 - 260,000
AI Hardware Reliability Engineer: Pod-Level RAS Telemetry
AI Hardware Reliability Engineer: Pod-Level RAS Telemetry

Intel • Santa Clara (CA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
On-site amenities
Engineering Manager, AI Kernel Reliability & Debugging
Engineering Manager, AI Kernel Reliability & Debugging

Cerebras • Raleigh (NC)

On-site
USD 180,000 - 260,000
Engineering Manager, Kernel Reliability
Engineering Manager, Kernel Reliability

Cerebras Systems • United States

On-site
USD 120,000 - 180,000
AI Hardware Kernel Performance Engineer
AI Hardware Kernel Performance Engineer

Cerebras • Sterling (VA)

On-site
USD 100,000 - 150,000
AI Hardware Systems Integration Engineer
AI Hardware Systems Integration Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Senior Systems Debug Engineer – High-Perf AI Hardware
Senior Systems Debug Engineer – High-Perf AI Hardware

Cerebras Systems • Sunnyvale (CA)

On-site
USD 190,000 - 230,000