Cluster Site Reliability Engineer

iFrame Corporation

Canada

On-site

CAD 294,000 - 476,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

iFrame Corporation in Canada is seeking a hands-on Senior Cluster SRE to own the physical reality of the platform in our seven regions. You will bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at SLA.

This is a hands-on role; you will see the hardware. Five-plus years operating large compute clusters; deep InfiniBand and Ethernet RoCE experience; comfortable writing Go or Python for tooling.

Qualifications

  • Five+ years operating large compute clusters (supercomputing centers, hyperscale infra, or HPC).
  • Deep InfiniBand and Ethernet RoCE experience: tuning, fabric debugging, lossless networking.
  • Comfort writing Go or Python for tooling; language not strictly required.

Responsibilities

  • Bring up new B200 / B300 / MI300X racks: cabling, ToR config, NCCL/RCCL validation, MFU baseline tests.
  • Drive InfiniBand fabric to spec (NDR/XDR) and chase residual bit-error budget to zero.
  • Run capacity planning across seven regions: forecast demand, model power/thermal headroom, coordinate with procurement on lead times.
  • Own the regional incident response: P1 incidents within 15 minutes; target four hours to resolve.
  • Build and maintain the bring-up runbook so the next hire can rack solo.
  • Carry the global pager about one week per six, plus runtime support.

Skills

Linux
InfiniBand
NCCL/RCCL
Kubernetes (host-level)
Terraform
Prometheus
Grafana
Go
Python

Tools

NVIDIA Bright / Base Command Manager
Tinkerbell
MAAS
Razor

Job description

Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware.

Type Full-time · IC4-IC6

Stack Linux InfiniBand (NDR / XDR) NCCL / RCCL Kubernetes (host-level) Terraform Prometheus / Grafana Go or Python

Hiring manager replies within 5 business days.

The team

About the team

Cluster SRE is six engineers across the seven regions. Each region has a primary and a secondary; you will be one of those for your region. The team coordinates daily, deploys weekly, and rotates a global pager.

Reports to the head of cluster engineering. Primary on a single region; rotates secondary for one neighboring region.

What you'll do
  • 01 Bring up new B200 / B300 / MI300X racks: cabling, ToR config, NCCL/RCCL all-reduce validation, MFU baseline tests.
  • 02 Drive InfiniBand fabric to spec - NDR / XDR depending on the rack - and chase residual bit-error budget down to zero.
  • 03 Run capacity planning across seven regions: forecast demand, model power and thermal headroom, work with procurement on lead times.
  • 04 Own the regional incident response. P1 incidents page within fifteen minutes; resolution target is four hours.
  • 05 Build and maintain the bring-up runbook so the second hire after you can do their first rack solo.
  • 06 Carry the global pager about one week per six, alongside runtime and customer engineering.

The bar

What we're looking for

Five-plus years operating large compute clusters — supercomputing centers, hyperscaler infra, or HPC at a national lab count.

Deep InfiniBand and Ethernet RoCE experience: subnet manager tuning, fabric debugging, lossless networking.

Comfort writing Go or Python for tooling. We are not strict about which.

Calm under load. You will be the named person on a $50M-ARR account when something goes wrong.

Nice to have, not required

Experience with NVIDIA Bright / Base Command Manager.

Bare-metal provisioning systems: Tinkerbell, MAAS, Razor, or in-house equivalents.

Compensation
In writing, like everything else

We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.

Base

$210,000 – $340,000 USD (US Cologix regions) / equivalent in CA.

Equity

Meaningful early-stage equity, refreshed on tenure milestones.

Notes

On-site pay differential at Cologix regions outside SF / NYC / Bay Area is +5-10% to compensate for travel.

Equal opportunity

We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU — bring it up on the manager call.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Site Reliability Engineer
Cluster Site Reliability Engineer

iFrame • Toronto

On-site
CAD 292,000 - 473,000
Regional Cluster SRE — GPU/HPC Infra Lead
Regional Cluster SRE — GPU/HPC Infra Lead

iFrame • Toronto

On-site
CAD 292,000 - 473,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
Software Engineer - GPU Fabric Observability
Software Engineer - GPU Fabric Observability

Baseten • Montreal (administrative region)

On-site
CAD 283,286 - 538,243
Equity
Health coverage
Flexible PTO
+4
Senior Product Manager (vMetal)
Senior Product Manager (vMetal)

vCluster Labs • Canada

On-site
USD 150,000 - 210,000
Competitive Salary
Equity
Platinum-Level Insurance
+2
Senior AI Infrastructure Engineer — HPC & Compute Clusters
Senior AI Infrastructure Engineer — HPC & Compute Clusters

Veeda AI • Toronto

On-site
CAD 100,000 - 150,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000