Systems Engineering Mechanical Staff, Reliability Engineer Toronto, Ontario, Canada

Tenstorrent

Toronto

On-site

CAD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tenstorrent is seeking a Staff Reliability Engineer to define reliability strategy and drive MTBF modeling across AI hardware platforms. You will partner with hardware, software, and manufacturing teams to ensure uptime, durability, and quality at scale.

The role demands 8+ years in reliability engineering, familiarity with HALT/HASS, and the ability to communicate risks clearly to leadership. This hybrid Toronto-based position offers opportunities across the hardware lifecycle.

Qualifications

  • 8+ years in reliability engineering (HPC, AI hardware, or data center systems preferred).
  • Experience leading root-cause investigations and driving fixes across engineering and supply chain.
  • Ability to summarize findings and feed back to Systems Engineering design team.

Responsibilities

  • Set reliability strategy for next-gen AI computing systems with MTBF modeling at core.
  • Lead root-cause investigations and drive cross-functional fixes.
  • Partner with System Dev & Compliance Validation to prepare hardware for testing and certification.
  • Travel to third-party test facilities for HALT/HASS and design risk assessment.
  • Collaborate with mechanical, electrical, thermal, software, and System Dev teams to align timelines.

Skills

Reliability engineering
MTBF modeling
Cross-functional leadership
Communication of risk

Education

Bachelor's or Master's in Mechanical, Electrical, or Reliability Engineering

Tools

HALT testing
HASS testing
Weibull analysis
FMEA

Job description

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.

Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI.

This role ishybrid, based out of Toronto, Canada.

We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.

Who You Are
  • You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems.
  • You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory.
  • You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership.
  • You're good at bringing people together, mechanical, electrical, thermal, software, and System Dev & Compliance Validation teams, especially when timelines are tight.
  • You hold a Bachelor's or Master's in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field.
What We Need
  • Someone to set the reliability strategy for our next-generation AI computing systems, with MTBF modeling as a core piece.
  • A strong problem-solver who can lead root-cause investigations and drive fixes across engineering and the supply chain.
  • Someone who summarizes findings and feeds them back to the Systems Engineering design team, reliability as an ongoing loop, not a one-time check.
  • A close partner to System Dev & Compliance Validation, helping set hardware up for success ahead of formal testing and certification.
  • Willingness to travel to third-party test facilities for HALT/HASS, to preplan, oversee testing, resolve DUT issues, and assess design risk in person.
What You Will Learn
  • How to build a reliability strategy from scratch for hardware that's pushing the boundaries of AI computing.
  • How to build predictive models, including MTBF frameworks and accelerated life testing.
  • How to turn test findings into design improvements through close collaboration with Systems Engineering.
  • How reliability work sets the stage for validation and certification success.
  • What it takes to run hands-on testing at manufacturing and test partner sites, including in Taiwan.
  • Tenstorrent offers a highly competitive compensation package and benefits

and we are an equal opportunity employer.

This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff, Reliability Engineer
Staff, Reliability Engineer

Tenstorrent • Toronto

On-site
CAD 140,000 - 200,000
Competitive compensation & benefits
System Hardware Technician
System Hardware Technician

Tenstorrent • Toronto

On-site
CAD 80,000 - 120,000
Competitive compensation
Benefits
Equal opportunity employer
Senior Engineer, System-Level Design Verification
Senior Engineer, System-Level Design Verification

Tenstorrent • Toronto

On-site
CAD 136,705 - 683,525
Competitive compensation and benefits
Sr. Engineer, Systems Debug
Sr. Engineer, Systems Debug

Tenstorrent • Toronto

On-site
CAD 90,000 - 140,000
Scale Out Software Engineer, Scale Out Toronto, Ontario, Canada
Scale Out Software Engineer, Scale Out Toronto, Ontario, Canada

Tenstorrent • Toronto

Hybrid
CAD 110,000 - 160,000
AI SW Software Engineer, Acceleration Kernel Development Toronto, Ontario, Canada
AI SW Software Engineer, Acceleration Kernel Development Toronto, Ontario, Canada

Tenstorrent • Toronto

On-site
CAD 100,000 - 500,000
DC Deployment Sr. Network Engineer Toronto, Ontario, Canada
DC Deployment Sr. Network Engineer Toronto, Ontario, Canada

Tenstorrent • Toronto

Hybrid
CAD 140,000 - 698,000
Competitive compensation
Equal opportunity employer
Hardware, Systems - Systems Engineering - Si Val / Qual Silicon Power & Characterization Engine[...]
Hardware, Systems - Systems Engineering - Si Val / Qual Silicon Power & Characterization Engine[...]

Tenstorrent Inc. • Toronto

On-site
CAD 139,000 - 700,000
Highly competitive compensation package
Benefits
Equal opportunity employer
Senior Embedded Engineer (AI IP)
Senior Embedded Engineer (AI IP)

Tenstorrent • Toronto

On-site
CAD 80,000 - 120,000
Silicon Power & Characterization Engineer
Silicon Power & Characterization Engineer

Tenstorrent • Toronto

On-site
CAD 100,000 - 500,000