Cluster Site Reliability Engineer

iFrame

Toronto

On-site

CAD 292,000 - 473,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.

You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.

Qualifications

  • Five-plus years operating large compute clusters in HPC or hyperscale environments.
  • Deep InfiniBand and Ethernet RoCE experience with fabric debugging.
  • Strong Linux fundamentals and kernel tuning knowledge.

Responsibilities

  • Bring up new racks, cabling, NCCL/RCCL tests.
  • Drive InfiniBand fabric to spec and reduce bit-error budget to zero.
  • Run capacity planning across seven regions and coordinate with procurement.
  • Own regional incident response with 4-hour resolution target.
  • Build and maintain bring-up runbook for future hires.
  • Manage global pager rotation and on-call duties.

Skills

Compute clusters
InfiniBand/Network tuning
Linux fundamentals
Go or Python
Calm under load

Tools

Kubernetes
Terraform
Prometheus
Grafana

Job description

Cluster Site Reliability Engineer

Cluster & SRE

Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware.

Team Cluster & SRE

Location: On-site

Region of choice: Cologix region of choice (Toronto, Montreal, Columbus, Vancouver, Ashburn)

Job Type and Tier

Full-time

  • IC4-IC6
Stack

Linux, InfiniBand (NDR / XDR), NCCL / RCCL, Kubernetes (host-level), Terraform, Prometheus / Grafana, Go or Python.

The team

Cluster SRE is six engineers across the seven regions. Each region has a primary and a secondary; you will be one of those for your region. The team coordinates daily, deploys weekly, and rotates a global pager.

Reports to the head of cluster engineering. Primary on a single region; rotates secondary for one neighboring region.

What you’ll do
  • 01 Bring up new B200 / B300 / MI300X racks: cabling, ToR config, NCCL/RCCL all-reduce validation, MFU baseline tests.
  • 02 Drive InfiniBand fabric to spec - NDR / XDR depending on the rack - and chase residual bit-error budget down to zero.
  • 03 Run capacity planning across seven regions: forecast demand, model power and thermal headroom, work with procurement on lead times.
  • 04 Own the regional incident response. P1 incidents page within fifteen minutes; resolution target is four hours.
  • 05 Build and maintain the bring-up runbook so the second hire after you can do their first rack solo.
  • 06 Carry the global pager about one week per six, alongside runtime and customer engineering.
What we’re looking for
  • Five-plus years operating large compute clusters - supercomputing centers, hyperscaler infra, or HPC at a national lab count.
  • Deep InfiniBand and Ethernet RoCE experience: subnet manager tuning, fabric debugging, lossless networking.
  • Strong Linux fundamentals: kernel parameters, NUMA topology, kernel cgroup limits.
  • Comfort writing Go or Python for tooling. We are not strict about which.
  • Calm under load. You will be the named person on a $50M-ARR account when something goes wrong.
Bonus (nice to have, not required)
  • DGX H100 / H200 / B200 bring-up history.
  • Experience with NVIDIA Bright / Base Command Manager.
  • Bare-metal provisioning systems: Tinkerbell, MAAS, Razor, or in-house equivalents.
  • Procurement / vendor management experience.
Compensation

In writing, like everything else

We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.

Base: $210,000 - $340,000 USD (US Cologix regions) / equivalent in CA.

Equity: Meaningful early-stage equity, refreshed on tenure milestones.

Notes: On-site pay differential at Cologix regions outside SF / NYC / Bay Area is +5-10% to compensate for travel.

Equal opportunity

We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU - bring it up on the manager call.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Site Reliability Engineer
Cluster Site Reliability Engineer

iFrame Corporation • Canada

On-site
CAD 294,000 - 476,000
Equity
Regional Cluster SRE — GPU/HPC Infra Lead
Regional Cluster SRE — GPU/HPC Infra Lead

iFrame • Toronto

On-site
CAD 292,000 - 473,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Site Reliability Engineering, Fabric (Mid, Senior, or Staff)
Site Reliability Engineering, Fabric (Mid, Senior, or Staff)

MongoDB • Toronto

Hybrid
CAD 144,000 - 200,000
Equity
Flexible paid time off
Parental leave
+2
Site Reliability Engineering, Fabric (Mid, Senior, or Staff)
Site Reliability Engineering, Fabric (Mid, Senior, or Staff)

MongoDB • Vancouver

Hybrid
CAD 144,000 - 200,000
Equity
Flexible paid time off
20 weeks fully-paid gender-neutral parental leave
+2
Site Reliability Engineer, Inference Infrastructure
Site Reliability Engineer, Inference Infrastructure

Visa Hunt • Toronto

Hybrid
CAD 120,000 - 180,000
Lunch stipend
Health & dental benefits
RRSP matching
+4
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000