Get more replies from employers
Send a job-specific resume in minutes.
iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.
You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.
iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.
You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.