Regional Cluster SRE — GPU/HPC Infra Lead

iFrame

Toronto

On-site

CAD 292,000 - 473,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.

You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.

Qualifications

  • Five-plus years operating large compute clusters in HPC or hyperscale environments.
  • Deep InfiniBand and Ethernet RoCE experience with fabric debugging.
  • Strong Linux fundamentals and kernel tuning knowledge.

Responsibilities

  • Bring up new racks, cabling, NCCL/RCCL tests.
  • Drive InfiniBand fabric to spec and reduce bit-error budget to zero.
  • Run capacity planning across seven regions and coordinate with procurement.
  • Own regional incident response with 4-hour resolution target.
  • Build and maintain bring-up runbook for future hires.
  • Manage global pager rotation and on-call duties.

Skills

Compute clusters
InfiniBand/Network tuning
Linux fundamentals
Go or Python
Calm under load

Tools

Kubernetes
Terraform
Prometheus
Grafana

Job description

iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.

You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI/ML HPC Infra & GPU Cluster
Senior SRE: AI/ML HPC Infra & GPU Cluster

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Cluster Site Reliability Engineer
Cluster Site Reliability Engineer

iFrame • Toronto

On-site
CAD 292,000 - 473,000
Cluster Site Reliability Engineer
Cluster Site Reliability Engineer

iFrame Corporation • Canada

On-site
CAD 294,000 - 476,000
Equity
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
GPU Infrastructure Operations & Technical Sales Engineer
GPU Infrastructure Operations & Technical Sales Engineer

Open People Network (OPN) • Toronto

Hybrid
CAD 90,000 - 130,000
Senior Site Reliability Engineer - Global Infra & CI/CD Impact
Senior Site Reliability Engineer - Global Infra & CI/CD Impact

CloudFactory Limited • Canada

Hybrid
CAD 120,000 - 160,000
Hybrid Working Model
Comprehensive medical cover
Group life insurance
+3
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior SRE — Distributed Data Platforms & Cloud Infra
Senior SRE — Distributed Data Platforms & Cloud Infra

OpenText • Southwestern Ontario

On-site
CAD 80,000 - 131,000
Senior SRE: Kubernetes Reliability for Cloud UI Services
Senior SRE: Kubernetes Reliability for Cloud UI Services

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training