ML Systems Integration Engineer

Cerebras

Toronto

On-site

CAD 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems is seeking a software engineer to participate in bringing up next‑generation AI hardware systems and supporting software. You will work across hardware and software domains to validate and stress distributed systems, while building automation that speeds debugging and validation tasks.

Ideal candidates have strong Python and C++, deep debugging skills, and a solid grounding in operating systems and computer architecture.

Qualifications

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related technical field.

Responsibilities

  • Participate in bring‑up of next‑generation AI hardware systems and supporting software infrastructure.
  • Debug complex system‑level issues spanning hardware and software interactions.
  • Investigate failures during system bring‑up and identify root causes using logs, telemetry, and diagnostic tools.
  • Build automation frameworks and internal tooling that improve system validation and debugging workflows.
  • Develop software used to test, validate, and stress distributed hardware systems during development and production cycles.
  • Collaborate closely with hardware engineers to isolate and resolve system integration issues.
  • Improve system observability by building tools that surface failures quickly and accelerate debugging.
  • Reproduce, triage, and diagnose difficult issues that arise during early hardware deployment.
  • Support validation and qualification of new hardware generations as systems move toward production readiness.
  • Continuously improve internal engineering workflows related to debugging, testing, and automation.

Skills

Python
C++
Debugging
OS fundamentals
Linux development
Computer architecture
Analytical thinking
Cross‑team collaboration
Communication

Education

BS or MS in Computer Science/Computer Engineering or Electrical Engineering

Tools

Automation frameworks
Internal tooling
Diagnostic tools
Telemetry/log analysis

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry‑leading training and inference speeds; over 10 times faster than GPU‑based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real‑time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting‑edge AI‑native startups. OpenAI recently announced a multi‑year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high‑speed inference.

Responsibilities
  • Participate in bring‑up of next‑generation AI hardware systems and supporting software infrastructure.
  • Debug complex system‑level issues spanning hardware and software interactions.
  • Investigate failures occurring during system bring‑up and identify root causes using logs, telemetry, and diagnostic tools.
  • Build automation frameworks and internal tooling that improve system validation and debugging workflows.
  • Develop software used to test, validate, and stress distributed hardware systems during development and production cycles.
  • Collaborate closely with hardware engineers to isolate and resolve system integration issues.
  • Improve system observability by building tools that surface failures quickly and accelerate debugging.
  • Reproduce, triage, and diagnose difficult issues that arise during early hardware deployment.
  • Support validation and qualification of new hardware generations as systems move toward production readiness.
  • Continuously improve internal engineering workflows related to debugging, testing, and automation.
Skills & Qualifications
  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related technical field.
  • Strong programming skills in Python and/or C++.
  • Excellent debugging and problem‑solving skills with ability to investigate complex technical issues methodically.
  • Solid understanding of operating systems fundamentals (processes, threads, memory management, concurrency, IPC).
  • Experience working in Linux development environments.
  • Understanding of computer architecture and interactions between hardware and software systems.
  • Strong analytical thinking and ability to break down complex system failures into actionable root causes.
  • Ability to work effectively across multiple engineering teams and collaborate in highly technical environments.
  • Strong communication skills and willingness to work on ambiguous technical problems.
Preferred Skills & Qualifications
  • Experience building automation frameworks, internal tooling, or test infrastructure
  • Familiarity with distributed systems concepts
  • Experience debugging large‑scale systems or complex infrastructure environments
  • Understanding of networking fundamentals and communication between distributed systems
  • Experience working with hardware‑adjacent software or system integration environments
  • Familiarity with performance analysis, system telemetry, and log analysis
  • Exposure to production systems validation or infrastructure reliability engineering

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third‑party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Software Engineer
Distributed Software Engineer

Cerebras Systems, Inc. • Ottawa

On-site
CAD 90,000 - 120,000
Job stability with startup vitality
Open access to cutting-edge AI research
ML Performance Benchmarking Engineer
ML Performance Benchmarking Engineer

Cerebras Systems, Inc. • Toronto

Hybrid
CAD 80,000 - 110,000
Job stability with startup vitality
Open-source cutting-edge AI research
Non-corporate work culture
ML Performance Benchmarking Engineer
ML Performance Benchmarking Engineer

Cerebras • Toronto

Hybrid
CAD 120,000 - 190,000
Groundbreaking technology platform
Equal opportunity employer
Innovative and collaborative work environment
CoDesign & NextGen Performance Engineer
CoDesign & NextGen Performance Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 180,000
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras Systems, Inc. • Toronto

On-site
CAD 140,000 - 210,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Toronto

On-site
CAD 150,000 - 210,000
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Software Engineer - New Grad 2026
Software Engineer - New Grad 2026

Cerebras Systems, Inc. • Toronto

Hybrid
CAD 70,000 - 90,000
Senior Runtime Engineer
Senior Runtime Engineer

Cerebras • Toronto

On-site
CAD 100,000 - 150,000
Non-corporate work culture
Stability with startup vitality
Opportunities for continuous learning