Get more replies from employers
Send a job-specific resume in minutes.
iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.
You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.
Cluster & SRE
Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware.
Location: On-site
Region of choice: Cologix region of choice (Toronto, Montreal, Columbus, Vancouver, Ashburn)
Full-time
Linux, InfiniBand (NDR / XDR), NCCL / RCCL, Kubernetes (host-level), Terraform, Prometheus / Grafana, Go or Python.
Cluster SRE is six engineers across the seven regions. Each region has a primary and a secondary; you will be one of those for your region. The team coordinates daily, deploys weekly, and rotates a global pager.
Reports to the head of cluster engineering. Primary on a single region; rotates secondary for one neighboring region.
In writing, like everything else
We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.
Base: $210,000 - $340,000 USD (US Cologix regions) / equivalent in CA.
Equity: Meaningful early-stage equity, refreshed on tenure milestones.
Notes: On-site pay differential at Cologix regions outside SF / NYC / Bay Area is +5-10% to compensate for travel.
We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU - bring it up on the manager call.