Full Stack LLM Engineer

Foundation Capital

Toronto

On-site

CAD 80,000 - 120,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary and benefits package
Opportunities for professional growth
Dynamic and innovative work environment

Job summary

A leading AI technology company in Toronto is seeking a versatile engineer to join their Inference Core Model Bringup team. This role involves bringing state-of-the-art models onto the Cerebras CSX systems, requiring experience in Python, compiler IRs, and strong debugging skills. Ideal candidates should have a degree in Computer Science or Engineering, and enjoy working in fast-paced environments. Benefits include a competitive salary and a chance to work with cutting-edge technologies.

Qualifications

  • Experience with deep learning frameworks such as PyTorch and TensorFlow.
  • Strong debugging skills necessary for performance and runtime integration.
  • Proficiency in C/C++ programming with emphasis on low-level optimization.

Responsibilities

  • Contribute to the end-to-end bring up of ML models on Cerebras CSX systems.
  • Work across model architecture translation, compiler optimizations, and performance tuning.
  • Debug performance and correctness issues related to models and hardware utilization.

Skills

Debugging skills
C/C++ programming
Python modeling code
Deep learning frameworks
Performance profiling

Education

Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field

Tools

LLVM
MLIR

Job description

About the Role

We are seeking a versatile and experienced engineer to join our Inference Core Model Bringup team. This team is responsible for rapidly bringing up state‑of‑the‑art open‑source models (like LLaMA, Qwen, etc.) or customer‑provided proprietary models on our Cerebras CSX systems. Success in this role requires a system‑minded generalist who thrives in fast‑paced bring‑up environments and is comfortable working across the entire Cerebras software stack. Your work will play a critical role in achieving unprecedented levels of performance, efficiency, and scalability for AI applications.

Responsibilities
  • Contribute to the end-to‑end bring up of ML models on Cerebras CSX systems.
  • Work across the stack: model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
  • Debug performance and correctness issues spanning model code, compiler IRs, runtime behavior, and hardware utilization.
  • Propose and prototype improvements across tools, APIs, or automation flows to accelerate future bring ups.
Skills & Qualifications
  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field.
  • Comfort navigating the full AI toolchain: Python modeling code, compiler IRs, performance profiling, etc.
  • Strong debugging skills across performance, numerical accuracy, and runtime integration.
  • Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and familiarity with model internals (e.g., attention, MoE, diffusion).
  • Proficiency in C/C++ programming and experience with low-level optimization.
  • Proven experience in compiler development, particularly with LLVM and/or MLIR.
  • Strong background in optimization techniques, particularly those involving NP‑hard problems.
What We Offer
  • Competitive salary and benefits package.
  • Opportunities for professional growth and career advancement.
  • A dynamic and innovative work environment.
  • The chance to work on cutting‑edge technologies and make a significant impact on the future of AI.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third‑party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Runtime and Kernel Engineer - Core ML
ML Runtime and Kernel Engineer - Core ML

Cerebras • Toronto

On-site
CAD 120,000 - 180,000
Staff Software Engineer, Inference API
Staff Software Engineer, Inference API

Engg • Toronto

On-site
CAD 140,000 - 200,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 110,000 - 160,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Foundation Capital • Toronto

On-site
CAD 120,000 - 180,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

Foundation Capital • Toronto

On-site
CAD 256,000 - 370,000
Sr. Staff Software Engineer, Inference Platform
Sr. Staff Software Engineer, Inference Platform

Cerebras • Toronto

On-site
CAD 140,000 - 210,000
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Foundation Capital • Toronto

On-site
CAD 110,000 - 170,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

Cerebras • Toronto

Hybrid
CAD 120,000 - 180,000
Inference ML API SDET
Inference ML API SDET

Foundation Capital • Toronto

Hybrid
CAD 199,000 - 299,000
Inference ML API SDET
Inference ML API SDET

Cerebras Systems • Lower Sackville

Hybrid
CAD 120,000 - 180,000