CoDesign & NextGen Performance Engineer

Foundation Capital

Toronto

On-site

CAD 110,000 - 170,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cerebras Systems builds an AI hardware platform designed to accelerate model training and inference at scale. The role focuses on characterizing, analyzing, and optimizing performance of state-of-the-art AI models running on Cerebras’ WSE hardware.

You will work across hardware and software to identify bottlenecks, improve efficiency, and influence next-generation architecture and software systems. Strong background in architecture and DL math is valued.

Qualifications

  • Bachelors / Masters / PhD in Electrical Engineering or Computer Science.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.
  • 3+ years of experience in a relevant domain (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
  • Exposure to performance profiling and debug on any system pipeline.
  • Comfort with C++ and Python.

Responsibilities

  • Bring up and optimize performance on new generations of the Cerebras WSE.
  • Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models.
  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.

Skills

C++
Python
Performance profiling
Kernel optimization
Computer architecture
HPC
CPU/GPU performance
Deep learning / LLM math

Education

Electrical Engineering or Computer Science degree

Tools

CPU/GPU simulators

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About The Role

This role focuses on characterizing, analyzing, and optimizing the performance of state-of-the-art AI models running on Cerebras’ breakthrough hardware. You will work across the hardware and software stack to identify bottlenecks, improve computational efficiency, and help influence the design of Cerebras’ next-generation AI architecture and software systems.

Responsibilities

  • Bring up and optimize performance on new generations of the Cerebras WSE.
  • Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models.
  • Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster.

Skills & Qualifications

  • Bachelors / Masters / PhD in Electrical Engineering or Computer Science. Strong background in computer architecture.
  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.
  • 3+ years of experience in a relevant domain (Computer Architecture, CPU/GPU Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
  • Exposure to performance profiling and debug on any system pipeline.
  • Comfort with C++ and Python.

Why Join Cerebras

  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 110,000 - 160,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Foundation Capital • Toronto

On-site
CAD 120,000 - 180,000
Co-Design & Next-Gen AI Performance Engineer
Co-Design & Next-Gen AI Performance Engineer

Foundation Capital • Toronto

On-site
CAD 110,000 - 170,000
FPGA Engineer
FPGA Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 160,000
Software Engineer - Tools & Infrastructure / DevOps
Software Engineer - Tools & Infrastructure / DevOps

Foundation Capital • Toronto

On-site
CAD 110,000 - 170,000
Senior Runtime Engineer
Senior Runtime Engineer

Cerebras • Toronto

On-site
CAD 100,000 - 150,000
Non-corporate work culture
Stability with startup vitality
Opportunities for continuous learning
FPGA Engineer
FPGA Engineer

Foundation Capital • Toronto

On-site
CAD 120,000 - 180,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

Cerebras • Toronto

Hybrid
CAD 120,000 - 180,000
Software Engineer - Tools & Infrastructure / DevOps
Software Engineer - Tools & Infrastructure / DevOps

Cerebras • Toronto

On-site
CAD 120,000 - 180,000
Senior Research Engineer - Inference ML
Senior Research Engineer - Inference ML

Cerebras Systems • Toronto

Hybrid
CAD 180,000 - 240,000