CoDesign & NextGen Performance Engineer

Cerebras Systems

Toronto

On-site

CAD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems is seeking a performance-focused engineer to characterize, analyze, and optimize AI models on the Wafer Scale Engine hardware. You will work across hardware and software to identify bottlenecks and improve efficiency, influencing Cerebras’ next-generation AI architecture and software systems.

Responsibilities include bringing up new hardware generations, building end-to-end performance models, and debugging runtime performance while developing visualization tools for performance

Qualifications

  • Bachelor’s or higher degree in Electrical Engineering or Computer Science.
  • Strong background in computer architecture.
  • Exposure to low-level deep learning / LLM math.
  • Analytical and problem-solving mindset.
  • 3+ years of related experience (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
  • Comfort with C++ and Python.

Responsibilities

  • Bring up and optimize performance on new generations of the Cerebras WSE.
  • Build performance models (kernel-level, end-to-end) to estimate performance of state-of-the-art models.
  • Optimize and debug kernel micro code and compiler algorithms to elevate inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to visualize performance data from the Wafer Scale Engine and compute cluster.

Education

Bachelors / Masters / PhD in Electrical Engineering or Computer Science
Strong background in computer architecture
Exposure to low-level deep learning / LLM math
Analytical and problem-solving mindset
3+ years of experience in Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC
Experience with CPU/GPU simulators
C++ and Python

Tools

CPU/GPU simulators

Job description

Cerebras Systems builds the world’s largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About The Role

This role focuses on characterizing, analyzing, and optimizing the performance of state-of-the-art AI models running on Cerebras’ breakthrough hardware. You will work across the hardware and software stack to identify bottlenecks, improve computational efficiency, and help influence the design of Cerebras’ next‑generation AI architecture and software systems.

Responsibilities
  • Bring up and optimize performance on new generations of the Cerebras WSE.
  • Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models.
  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.
Skills & Qualifications
  • Bachelors / Masters / PhD in Electrical Engineering or Computer Science.
    Strong background in computer architecture.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.
  • 3+ years of experience in a relevant domain (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
  • Exposure to performance profiling and debug on any system pipeline.
  • Comfort with C++ and Python.
Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third‑party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Benchmarking Engineer
ML Performance Benchmarking Engineer

Cerebras Systems, Inc. • Toronto

Hybrid
CAD 80,000 - 110,000
Job stability with startup vitality
Open-source cutting-edge AI research
Non-corporate work culture
Distributed Software Engineer
Distributed Software Engineer

Cerebras Systems, Inc. • Ottawa

On-site
CAD 90,000 - 120,000
Job stability with startup vitality
Open access to cutting-edge AI research
ML Performance Benchmarking Engineer
ML Performance Benchmarking Engineer

Cerebras • Toronto

Hybrid
CAD 120,000 - 190,000
Groundbreaking technology platform
Equal opportunity employer
Innovative and collaborative work environment
Senior Runtime Engineer
Senior Runtime Engineer

Cerebras • Toronto

On-site
CAD 100,000 - 150,000
Non-corporate work culture
Stability with startup vitality
Opportunities for continuous learning
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 90,000 - 150,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Toronto

On-site
CAD 150,000 - 210,000
Simulation Engineer
Simulation Engineer

Cerebras Systems • Toronto

On-site
CAD 70,000 - 90,000
Simulation Engineer
Simulation Engineer

Cerebras • Toronto

On-site
CAD 90,000 - 120,000
Host and Network IO FPGA Engineer
Host and Network IO FPGA Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 180,000
DevOps Engineer - New Grad 2026
DevOps Engineer - New Grad 2026

Cerebras Systems, Inc. • Toronto

On-site
CAD 70,000 - 90,000
Opportunity to work on an innovative AI platform
Diverse and inclusive work environment
Job stability with startup vitality