ML Systems Performance Engineer

Cerebras

United States

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras is seeking engineers for the inference performance team, focusing on optimizing model inference speed and throughput. Responsibilities include building performance models, debugging kernel micro code, and developing tools for visualizing performance data.

The ideal candidate holds a degree in Electrical Engineering or Computer Science and has 3+ years of experience in CPU/GPU performance and kernel optimization. A diverse, equal opportunity environment is promoted at Cerebras, fostering inclusive teams.

Qualifications

  • 3+ years of experience in Computer Architecture, CPU/GPU Performance, or Kernel Optimization.
  • Strong analytical and problem-solving mindset.
  • Understanding of low-level deep learning / LLM math.

Responsibilities

  • Build performance models to estimate performance of ML models.
  • Optimize and debug kernel micro code and compiler algorithms.
  • Develop tools to visualize performance data.

Skills

Computer architecture
C++
Python
Kernel optimization
Performance profiling

Education

Bachelors / Masters / PhD in Electrical Engineering or Computer Science

Tools

CPU/GPU simulators

Job description

About The Role

Engineers on the inference performance team operate at the intersection of hardware and software, driving end-to-end model inference speed and throughput. Their work spans low-level kernel performance debugging and optimization, system-level performance analysis, performance modeling and estimation, and the development of tooling for performance projection and diagnostics.

Responsibilities
  • Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models.
  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.
Requirements
  • Bachelors / Masters / PhD in Electrical Engineering or Computer Science.
  • Strong background in computer architecture.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.
  • 3+ years of experience in a relevant domain (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
  • Exposure to performance profiling and debug on any system pipeline.
  • Comfort with C++ and Python.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras Systems • United States

On-site
USD 110,000 - 150,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
ML Inference Performance Architect
ML Inference Performance Architect

Cerebras • United States

On-site
USD 100,000 - 130,000
ML Inference Performance Engineer - Kernel Optimizer
ML Inference Performance Engineer - Kernel Optimizer

Cerebras Systems • United States

On-site
USD 110,000 - 150,000
Staff Kernel Optimzation Engineer
Staff Kernel Optimzation Engineer

Cerebras • Sterling (VA)

On-site
USD 100,000 - 150,000
Kernel Engineer
Kernel Engineer

Cerebras • Raleigh (NC)

On-site
USD 100,000 - 140,000
Equal opportunity work environment
Continuous learning and support
Diverse team culture
Kernel Engineer - New Grad
Kernel Engineer - New Grad

Cerebras • United States

On-site
USD 120,000 - 180,000
Senior Performance Engineer, Inference
Senior Performance Engineer, Inference

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Opportunities for contributing to open-source projects
Stable work environment with startup vitality
Kernel Engineer
Kernel Engineer

Cerebras • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture