An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Get past ATS filters
Job summary
A technology company in Toronto seeks a Senior Site Reliability Engineer to manage and optimize its HPC infrastructure. In this role, you'll ensure smooth operations of a powerful GPU cluster, deploy infrastructure-as-code solutions, and support ML teams. Candidates should have extensive SRE experience, proficiency in Linux, and familiarity with Kubernetes and Ceph storage. This position offers the chance to work with cutting-edge technology in a collaborative environment, perfect for problem-solvers who love learning.
Qualifications
5+ years of experience in Site Reliability Engineering or HPC operations.
Proficiency in Linux systems administration, specifically Ubuntu/Debian.
Experience with Kubernetes and container orchestration.
Knowledge of security best practices in multi-tenant environments.
Strong grasp of networking fundamentals, specifically L2/L3.
Responsibilities
Manage and optimize operations of the HPC cluster.
Deploy and maintain infrastructure-as-code solutions.
Support ML/research teams in optimizing cluster usage.
Troubleshoot and optimize Ceph storage clusters.
Develop automation and tooling for efficiency.
Skills
Linux systems administration
Kubernetes
Python scripting
Bash scripting
Ceph storage management
Tools
Ansible
Terraform
Job description
A technology company in Toronto seeks a Senior Site Reliability Engineer to manage and optimize its HPC infrastructure. In this role, you'll ensure smooth operations of a powerful GPU cluster, deploy infrastructure-as-code solutions, and support ML teams. Candidates should have extensive SRE experience, proficiency in Linux, and familiarity with Kubernetes and Ceph storage. This position offers the chance to work with cutting-edge technology in a collaborative environment, perfect for problem-solvers who love learning.