Hardware Analytics Engineer

Cerebras

Sunnyvale (CA)

On-site

USD 213,675 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Telecommuting permitted

Job summary

Cerebras Systems Inc. in Sunnyvale, CA is seeking a Hardware Analytics Engineer to design and operate scalable data pipelines for telemetry, reliability analytics, and performance optimization across AI server platforms.

You will develop frameworks using Python, SQL, Tableau, Hive, and Spark to forecast hardware failures, lead thermal studies, and deliver actionable recommendations to improve efficiency and sustainability.

Qualifications

  • Master’s degree or foreign equivalent in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
  • 3 years of experience as Hardware Analytics Engineer or related roles.

Responsibilities

  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams across heterogeneous compute, storage, and AI server platforms.
  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.
  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.
  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.

Skills

Data pipelines
ETL
Python
SQL
Tableau
Linux
Automation
ML for hardware
Anomaly detection
Reliability analytics
Performance optimization

Education

Master's degree or foreign equivalent in Electrical Engineering, Computer Engineering, Computer Science, or related field

Tools

Hive
Spark
Tableau

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

Cerebras Systems Inc. has multiple openings for Hardware Analytics Engineer

Title: Hardware Analytics Engineer

Job Duties:
  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.
  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.
  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.
  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.
Minimum Requirements:

Master’s degree or foreign equivalent degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field and 3 years of experience as Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or a related occupation required.

Required Skills:
  • Large-scale data pipeline architecture and ETL, distributed data processing (Hive, Spark), and dashboard development;
  • Python, SQL, Tableau, Linux, and automation scripting;
  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction;
  • Predictive modeling, statistical analysis, A/B testing, anomaly detection, and data visualization in hardware reliability and performance; and
  • Hardware analytics for compute, storage, and AI servers; power and thermal optimization; GPU burn-in efficiency optimization; and reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD.
Additional Information:

Employer’s name: Cerebras Systems Inc.

Job site : 1237 E Arques Avenue, Sunnyvale, CA 94085

Telecommuting permitted

Salary Range: $213,675.00 per year to $225,000.00 per year

Why Join Cerebras
  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Hardware Technical Program Manager
Senior Hardware Technical Program Manager

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 180,000 - 230,000
Working on cutting-edge AI technology
Inclusive and diverse work environment
Opportunity for continuous learning and growth
ML Software Tool Development Engineer
ML Software Tool Development Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Equal opportunity employer
Non-corporate work culture
Continuous learning opportunities
Mechanical Engineer
Mechanical Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 180,000 - 200,000
Job stability with startup vitality
Non-corporate work culture
Opportunity to work on leading AI technologies
ML Software Tool Development Engineer
ML Software Tool Development Engineer

Cerebras • Raleigh (NC)

On-site
USD 140,000 - 210,000
Senior Hardware Technical Program Manager
Senior Hardware Technical Program Manager

Cerebras Systems • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Innovative work culture
Opportunity for groundbreaking research
Job stability with startup vitality
Hardware Analytics Engineer: AI Telemetry & Reliability
Hardware Analytics Engineer: AI Telemetry & Reliability

Cerebras • United States

Remote
USD 120,000 - 190,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
Infrastructure Engineer (Data Center Operations)
Infrastructure Engineer (Data Center Operations)

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 170,000
Infrastructure Engineer (Data Center Operations)
Infrastructure Engineer (Data Center Operations)

Socket.dev • Sunnyvale (CA)

On-site
USD 120,000 - 190,000
Mechanical Engineer
Mechanical Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 200,000