Senior SRE: AI/ML HPC Infra & GPU Cluster

Boson AI

Toronto

On-site

CAD 100,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A technology company in Toronto seeks a Senior Site Reliability Engineer to manage and optimize its HPC infrastructure. In this role, you'll ensure smooth operations of a powerful GPU cluster, deploy infrastructure-as-code solutions, and support ML teams. Candidates should have extensive SRE experience, proficiency in Linux, and familiarity with Kubernetes and Ceph storage. This position offers the chance to work with cutting-edge technology in a collaborative environment, perfect for problem-solvers who love learning.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or HPC operations.
  • Proficiency in Linux systems administration, specifically Ubuntu/Debian.
  • Experience with Kubernetes and container orchestration.
  • Knowledge of security best practices in multi-tenant environments.
  • Strong grasp of networking fundamentals, specifically L2/L3.

Responsibilities

  • Manage and optimize operations of the HPC cluster.
  • Deploy and maintain infrastructure-as-code solutions.
  • Support ML/research teams in optimizing cluster usage.
  • Troubleshoot and optimize Ceph storage clusters.
  • Develop automation and tooling for efficiency.

Skills

Linux systems administration
Kubernetes
Python scripting
Bash scripting
Ceph storage management

Tools

Ansible
Terraform

Job description

A technology company in Toronto seeks a Senior Site Reliability Engineer to manage and optimize its HPC infrastructure. In this role, you'll ensure smooth operations of a powerful GPU cluster, deploy infrastructure-as-code solutions, and support ML teams. Candidates should have extensive SRE experience, proficiency in Linux, and familiarity with Kubernetes and Ceph storage. This position offers the chance to work with cutting-edge technology in a collaborative environment, perfect for problem-solvers who love learning.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
AI/ML Infrastructure Engineer — SRE for Scalable AI Clusters
AI/ML Infrastructure Engineer — SRE for Scalable AI Clusters

BULL-IT SOLUTIONS LTD • Montreal

On-site
CAD 100,000 - 130,000
Senior SRE: Linux & Cloud Infrastructure
Senior SRE: Linux & Cloud Infrastructure

Atlantis IT Group • Montreal

On-site
CAD 80,000 - 100,000
Senior SRE Leader: Scale Reliability & Observability
Senior SRE Leader: Scale Reliability & Observability

Rootly • Toronto

On-site
CAD 120,000 - 180,000
Competitive compensation
Comprehensive medical coverage
3 weeks of vacation
+2
Senior SRE: Kubernetes Reliability for Cloud UI Services
Senior SRE: Kubernetes Reliability for Cloud UI Services

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Senior SRE: Data Center Networking & Reliability
Senior SRE: Data Center Networking & Reliability

Google • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Equity
Bonus target
Benefits
Senior SRE: Automation, Observability & Batch Performance
Senior SRE: Automation, Observability & Batch Performance

Kyndryl • Toronto

Hybrid
CAD 100,000 - 130,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
Senior SRE/DevOps Engineer - Kubernetes & Observability
Senior SRE/DevOps Engineer - Kubernetes & Observability

Infotek Consulting Inc. • Toronto

Hybrid
CAD 125,000 - 150,000
GPU Infrastructure Engineer for AI Inference & Serving
GPU Infrastructure Engineer for AI Inference & Serving

P2P • Montreal (administrative region)

On-site
CAD 90,000 - 120,000