ML Inference Performance Architect

Cerebras

United States

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cerebras is seeking engineers for the inference performance team, focusing on optimizing model inference speed and throughput. Responsibilities include building performance models, debugging kernel micro code, and developing tools for visualizing performance data.

The ideal candidate holds a degree in Electrical Engineering or Computer Science and has 3+ years of experience in CPU/GPU performance and kernel optimization. A diverse, equal opportunity environment is promoted at Cerebras, fostering inclusive teams.

Qualifications

  • 3+ years of experience in Computer Architecture, CPU/GPU Performance, or Kernel Optimization.
  • Strong analytical and problem-solving mindset.
  • Understanding of low-level deep learning / LLM math.

Responsibilities

  • Build performance models to estimate performance of ML models.
  • Optimize and debug kernel micro code and compiler algorithms.
  • Develop tools to visualize performance data.

Skills

Computer architecture
C++
Python
Kernel optimization
Performance profiling

Education

Bachelors / Masters / PhD in Electrical Engineering or Computer Science

Tools

CPU/GPU simulators

Job description

Cerebras is seeking engineers for the inference performance team, focusing on optimizing model inference speed and throughput. Responsibilities include building performance models, debugging kernel micro code, and developing tools for visualizing performance data.

The ideal candidate holds a degree in Electrical Engineering or Computer Science and has 3+ years of experience in CPU/GPU performance and kernel optimization. A diverse, equal opportunity environment is promoted at Cerebras, fostering inclusive teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras • United States

On-site
USD 100,000 - 130,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
AI Inference Performance Engineer
AI Inference Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Advanced Technology: AI/ML Research Scientist
Advanced Technology: AI/ML Research Scientist

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
Kernel Engineer
Kernel Engineer

Cerebras • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture