HPC Engineer

Wyatt Partners

Toronto

On-site

CAD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading company in AI technologies is seeking an experienced HPC engineer to manage large AI clusters with Nvidia chips. This position involves leading a team and ensuring efficient operations while encouraging applications from individuals looking to further their managerial experience.

Qualifications

  • Experience as an HPC engineer.
  • Familiarity with scheduler technology and debugging.
  • Ability to discuss past cluster issues and learning experiences.

Responsibilities

  • Lead a team responsible for maintaining AI clusters.
  • Ensure the smooth operation of major AI clusters.

Skills

HPC engineering
Scheduler technology
Debugging

Job description

If you are interested in managing large AI clusters equipped with Nvidia AI chips, this role could be a great fit for you.

You will lead a team and be responsible for maintaining the environment and ensuring the smooth operation of one of our client's major AI Clusters.

We welcome applications from engineers who have not previously led a team or see this role as a step up in their responsibilities.

Essentially, if you have experience as an HPC engineer, familiarity with Scheduler technology and debugging, and can discuss past experiences where things went wrong in a cluster and what you learned from those situations, we would love to hear from you.

We are also in the process of expanding the wider team under this individual, so we are open to conversations across the board.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — HPC & Compute Clusters
Senior AI Infrastructure Engineer — HPC & Compute Clusters

Veeda AI • Toronto

On-site
CAD 100,000 - 150,000
Senior AI Compute Cluster Operations Engineer
Senior AI Compute Cluster Operations Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
HPC Specialist
HPC Specialist

Arcadion • Ottawa

On-site
CAD 80,000 - 110,000
Senior AI Cluster Operations Engineer
Senior AI Cluster Operations Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Staff / Principal Software Engineer
Staff / Principal Software Engineer

United States Digital Space LLC • Toronto

On-site
CAD 140,000 - 210,000
Staff / Principal Software Engineer
Staff / Principal Software Engineer

High Tech Genesis Inc. • Toronto

On-site
CAD 140,000 - 190,000