ML Systems Performance Engineer

Cerebras Systems

United States

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems is seeking engineers for the inference performance team, focusing on optimizing AI model inference speed on the world’s largest AI chip. This role involves performance modeling, debugging, and tooling development to enhance machine learning capabilities.

Ideal candidates should have a relevant degree, strong experience in computer architecture, and proficiency in C++ and Python. Join us at Cerebras to help shape the future of AI technology.

Qualifications

  • Strong background in computer architecture.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • 3+ years of experience in a relevant domain.

Responsibilities

  • Build performance models to estimate ML model performance.
  • Optimize and debug kernel micro code for higher inference speed.
  • Develop tools to visualize performance data.

Skills

Kernel optimization
C++
Python
Performance profiling
Problem-solving

Education

Bachelors / Masters / PhD in Electrical Engineering or Computer Science

Tools

CPU/GPU simulators

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry‑leading training and inference speeds; over 10 times faster than GPU‑based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real‑time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting‑edge AI‑native startups. OpenAI recently announced a multi‑year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high‑speed inference.

About The Role

Engineers on the inference performance team operate at the intersection of hardware and software, driving end‑to‑end model inference speed and throughput. Their work spans low‑level kernel performance debugging and optimization, system‑level performance analysis, performance modeling and estimation, and the development of tooling for performance projection and diagnostics.

Responsibilities
  • Build performance models (kernel‑level, end‑to‑end) to estimate the performance of state of the art and customer ML models.

  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.

  • Debug and understand runtime performance on the system and cluster.

  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.

Requirements
  • Bachelors / Masters / PhD in Electrical Engineering or Computer Science.

  • Strong background in computer architecture.

  • Exposure to and understanding of low‑level deep learning / LLM math.

  • Strong analytical and problem‑solving mindset.

  • 3+ years of experience in a relevant domain (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).

  • Experience working on CPU/GPU simulators.

  • Exposure to performance profiling and debug on any system pipeline.

  • Comfort with C++ and Python.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around us.

This website or its third‑party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras • United States

On-site
USD 100,000 - 130,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Advanced Technology: AI/ML Research Scientist
Advanced Technology: AI/ML Research Scientist

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Kernel Engineer
Kernel Engineer

Cerebras • Raleigh (NC)

On-site
USD 100,000 - 140,000
Equal opportunity work environment
Continuous learning and support
Diverse team culture
Kernel Engineer
Kernel Engineer

Cerebras Systems • United States

On-site
USD 110,000 - 140,000
Non-corporate work culture
Equal opportunity employer
Continuous learning and growth opportunities
Kernel Engineer
Kernel Engineer

Cerebras • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture
ML Software Tool Development Engineer
ML Software Tool Development Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Equal opportunity employer
Non-corporate work culture
Continuous learning opportunities
Senior Performance Engineer, Inference
Senior Performance Engineer, Inference

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Opportunities for contributing to open-source projects
Stable work environment with startup vitality
ML Software Tool Development Engineer
ML Software Tool Development Engineer

Cerebras • Raleigh (NC)

On-site
USD 140,000 - 210,000