Remote Hardware Analytics Engineer: AI Server Telemetry

Cerebras

Sunnyvale (CA)

On-site

USD 213,675 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Telecommuting permitted

Job summary

Cerebras Systems Inc. in Sunnyvale, CA is seeking a Hardware Analytics Engineer to design and operate scalable data pipelines for telemetry, reliability analytics, and performance optimization across AI server platforms.

You will develop frameworks using Python, SQL, Tableau, Hive, and Spark to forecast hardware failures, lead thermal studies, and deliver actionable recommendations to improve efficiency and sustainability.

Qualifications

  • Master’s degree or foreign equivalent in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
  • 3 years of experience as Hardware Analytics Engineer or related roles.

Responsibilities

  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams across heterogeneous compute, storage, and AI server platforms.
  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.
  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.
  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.

Skills

Data pipelines
ETL
Python
SQL
Tableau
Linux
Automation
ML for hardware
Anomaly detection
Reliability analytics
Performance optimization

Education

Master's degree or foreign equivalent in Electrical Engineering, Computer Engineering, Computer Science, or related field

Tools

Hive
Spark
Tableau

Job description

Cerebras Systems Inc. in Sunnyvale, CA is seeking a Hardware Analytics Engineer to design and operate scalable data pipelines for telemetry, reliability analytics, and performance optimization across AI server platforms.

You will develop frameworks using Python, SQL, Tableau, Hive, and Spark to forecast hardware failures, lead thermal studies, and deliver actionable recommendations to improve efficiency and sustainability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Analytics Engineer: AI Telemetry & Reliability
Hardware Analytics Engineer: AI Telemetry & Reliability

Cerebras • United States

Remote
USD 120,000 - 190,000
Hardware Analytics Engineer
Hardware Analytics Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 213,000 - 225,000
Telecommuting permitted
AI Hardware Systems Integration Engineer
AI Hardware Systems Integration Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Senior Systems Debug Engineer – High-Perf AI Hardware
Senior Systems Debug Engineer – High-Perf AI Hardware

Cerebras Systems • Sunnyvale (CA)

On-site
USD 190,000 - 230,000
Senior Mechanical Engineer - AI Data Center Delivery
Senior Mechanical Engineer - AI Data Center Delivery

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
AI Hardware Reliability Engineer: Pod-Level RAS Telemetry
AI Hardware Reliability Engineer: Pod-Level RAS Telemetry

Intel • Santa Clara (CA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
On-site amenities
AI Hardware Kernel Performance Engineer
AI Hardware Kernel Performance Engineer

Cerebras • Sterling (VA)

On-site
USD 100,000 - 150,000
Staff AI Telemetry Architect for High-Scale Observability
Staff AI Telemetry Architect for High-Scale Observability

Bitdeer • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Next-Gen AI Performance Engineer — Optimize Breakthrough Hardware
Next-Gen AI Performance Engineer — Optimize Breakthrough Hardware

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Senior AI Inference Reliability Engineer
Senior AI Inference Reliability Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Inclusive work environment
Opportunities for continuous learning
Startup vitality with job stability