Site Reliability Engineer, AI/ML Infrastructure

Boson AI

Toronto

On-site

CAD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company in Toronto seeks a Senior Site Reliability Engineer to manage and optimize its HPC infrastructure. In this role, you'll ensure smooth operations of a powerful GPU cluster, deploy infrastructure-as-code solutions, and support ML teams. Candidates should have extensive SRE experience, proficiency in Linux, and familiarity with Kubernetes and Ceph storage. This position offers the chance to work with cutting-edge technology in a collaborative environment, perfect for problem-solvers who love learning.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or HPC operations.
  • Proficiency in Linux systems administration, specifically Ubuntu/Debian.
  • Experience with Kubernetes and container orchestration.
  • Knowledge of security best practices in multi-tenant environments.
  • Strong grasp of networking fundamentals, specifically L2/L3.

Responsibilities

  • Manage and optimize operations of the HPC cluster.
  • Deploy and maintain infrastructure-as-code solutions.
  • Support ML/research teams in optimizing cluster usage.
  • Troubleshoot and optimize Ceph storage clusters.
  • Develop automation and tooling for efficiency.

Skills

Linux systems administration
Kubernetes
Python scripting
Bash scripting
Ceph storage management

Tools

Ansible
Terraform

Job description

Overview

We2;re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers.

Youll be hands-on with the full lifecycle of HPC infrastructure: planning, building, testing, deploying, and keeping everything running smoothly. That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working closely with engineering and science teams to ensure they have what they need. Youll also help us plan for future capacity and evaluate new technologies as we continue to scale.

Responsibilities
  • Manage and optimize HPC cluster operations
  • Deploy and maintain infrastructure-as-code solutions
  • Support ML/research teams with cluster usage optimization
  • Operate, troubleshoot and optimize Ceph storage clusters
  • Develop automation and tooling
Minimum Qualifications
  • 5+ years of experience in SRE or HPC operations
  • Proficiency in Linux systems administration (Ubuntu/Debian)
  • Experience with Kubernetes and container orchestration
  • Experience with Ceph >1PB deployments and maintenance
  • Knowledge of security best practices in multi-tenant environments
  • Understanding of L2/L3 networking fundamentals
  • Skilled in Python and Bash scripting
Preferred Qualifications
  • Experience with infrastructure-as-code tools (Ansible/Terraform)
  • Experience with GitOps (Helm, ArgoCD)
  • Strong grasp of RDMA, InfiniBand, and GPUDirect technologies
  • Familiarity with deep learning frameworks such as PyTorch and TensorFlow
  • Familiarity in at least one cloud platform: AWS, Azure or GCP

If youre a natural problem-solver with a passion for continuous learning, wed love to hear from you.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI/ML HPC Infra & GPU Cluster
Senior SRE: AI/ML HPC Infra & GPU Cluster

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Senior AI Infrastructure Engineer — HPC & Compute Clusters
Senior AI Infrastructure Engineer — HPC & Compute Clusters

Veeda AI • Toronto

On-site
CAD 100,000 - 150,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
AI/ML Infrastructure Engineer
AI/ML Infrastructure Engineer

BULL-IT SOLUTIONS LTD • Montreal

On-site
CAD 100,000 - 130,000
HPC Specialist
HPC Specialist

DRW Holdings, LLC. • Montreal (administrative region)

On-site
CAD 100,000 - 130,000
Cluster Site Reliability Engineer
Cluster Site Reliability Engineer

iFrame • Toronto

On-site
CAD 292,000 - 473,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Senior Engineer-Cloud AI Infrastructure
Senior Engineer-Cloud AI Infrastructure

Huawei Canada • Markham

On-site
CAD 172,000 - 306,000