System Administrator

Zohorecruit

Ottawa, Outaouais

Sur place

CAD 90 000 - 120 000

Plein temps

14 jours+
Générateur de candidature

Une candidature conçue pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.

Passez les filtres ATS

Résumé du poste

City Government of Canada in Ottawa seeks an experienced System Administrator, Level 2, to manage HPC clusters and support scientists who rely on them. You will install, run, and debug applications, ensuring researchers stay focused on science rather than IT issues.

Responsibilities include maintaining clusters, diagnosing incidents, and collaborating with scientists to meet their HPC needs. Strong Linux, schedulers, CUDA, and MPI skills required.

Qualifications

  • Solid Linux HPC cluster administration experience.
  • Hands-on experience with job schedulers such as PBS Pro/Torque, SLURM, or SGE.
  • Experience troubleshooting CUDA installations, GPU failures, and driver issues.
  • Familiarity with scientific computing toolchains — compilers (GNU, Intel), MPI implementations, EasyBuild, and Spack.
  • Experience supporting researchers or end-users with application builds and runtime issues.
  • Working knowledge of configuration management tools (Git, Ansible, MS DevOps).
  • Comfortable working independently and producing clear technical documentation.
  • Eligible to obtain and maintain a Secret-level security clearance.

Responsabilités

  • Maintain the HPC cluster — hardware, image management, local networking, scheduler, and backups.
  • Troubleshoot environment incidents to ensure a quick return to normal operations.
  • Meet with scientists to evaluate their HPC support requirements.
  • Develop task plans to meet researchers' needs, consulting the technical authority for approval.
  • Support application builds, installs, and runtime troubleshooting (GNU, Intel, Fortran, Nvidia).
  • Support open-source and commercial software, including Python/Anaconda installs, Bash scripting, build/make tools, EasyBuild, Spack, and MPI implementations (MPICH, OpenMPI, IntelMPI, HPMPI).
  • Assist with compilation and runtime of in-house developed applications.
  • General systems management.
  • Manage Linux OS patching schedules and reliability.
  • Manage user accounts (creation, deletion) and environment modules.
  • Manage configuration via Git, MS DevOps, and Ansible Playbooks.
  • Manage RPM/DEB packages and troubleshoot ThinLinc.
  • Troubleshoot jobs on schedulers (PBS Pro/Torque, SLURM, SGE).
  • Ensure reliable CUDA installs; troubleshoot GPU failures and CUDA software/driver issues.
  • Provide hardware support — memory upgrades, storage arrays, power/network cabling, ILO.
  • Documentation.

Connaissances

Linux HPC
PBS Pro/Torque
SLURM
SGE
CUDA troubleshooting
GPU drivers
MPI
EasyBuild
Spack
Python scripting
Bash scripting
Git
Ansible
MS DevOps
Documentation

Outils

PBS Pro/Torque
SLURM
SGE
CUDA
MPI (MPICH/OpenMPI/IntelMPI)
EasyBuild
Spack
Git
Ansible
MS DevOps

Description du poste

  • City Government of Canada Ottawa and Gatineau offices
  • Country Canada
Job Description

MGIS is seeking a System Administrator, Level 2, to manage High Performance Computing (HPC) clusters and support the scientists who rely on them. This role blends HPC system administration with hands-on user support — helping researchers install, run, and debug applications on HPC infrastructure so they can focus on their science instead of IT issues.

HPC environments in scope include clustered CPU/GPU systems with job schedulers and attached parallel storage (e.g., Lustre, GPFS).

What you'll be doing

Maintain the HPC cluster — hardware, image management, local networking, scheduler, and backups

Troubleshoot environment incidents to ensure a quick return to normal operations

Meet with scientists to evaluate their HPC support requirements

Develop task plans to meet researchers' needs, consulting the technical authority for approval

Support application builds, installs, and runtime troubleshooting (GNU, Intel, Fortran, Nvidia)

Support open-source and commercial software, including Python/Anaconda installs, Bash scripting, build/make tools, EasyBuild, Spack, and MPI implementations (MPICH, OpenMPI, IntelMPI, HPMPI)

Assist with compilation and runtime of in-house developed applications

General systems management

Manage Linux OS patching schedules and reliability

Manage user accounts (creation, deletion) and environment modules

Manage configuration via Git, MS DevOps, and Ansible Playbooks

Manage RPM/DEB packages and troubleshoot ThinLinc

Troubleshoot jobs on schedulers (PBS Pro/Torque, SLURM, SGE)

Ensure reliable CUDA installs; troubleshoot GPU failures and CUDA software/driver issues

Provide hardware support — memory upgrades, storage arrays, power/network cabling, ILO

Documentation

Document every process and task to support enterprise knowledge continuity

Submit weekly progress reports to the Technical Authority

Requirements

What we're looking for

Solid experience administering Linux-based HPC clusters (CPU/GPU nodes, schedulers, parallel storage)

Hands-on experience with job schedulers such as PBS Pro/Torque, SLURM, or SGE

Experience troubleshooting CUDA installations, GPU failures, and driver issues

Familiarity with scientific computing toolchains — compilers (GNU, Intel), MPI implementations, EasyBuild, and Spack

Experience supporting researchers or end-users with application builds and runtime issues

Working knowledge of configuration management tools (Git, Ansible, MS DevOps)

Comfortable working independently and producing clear technical documentation

Eligible to obtain and maintain a Secret-level security clearance

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

System Administrator
System Administrator

MGIS Inc. • Canada

Sur place
CAD 80 000 - 100 000
Cloud Sales Representative (CSR), AWS Canada
Cloud Sales Representative (CSR), AWS Canada

amazon web services canada • Vancouver

Sur place
CAD 171 000 - 256 000
Datacenter Technician
Datacenter Technician

Boson AI • Barrie

Sur place
CAD 50 000 - 100 000
Shared Services - Systems Administrator P2
Shared Services - Systems Administrator P2

Osimaritime • Burnaby

Sur place
CAD 65 000 - 95 000
System Administrator
System Administrator

Hexagon's Autonomy & Positioning division • Calgary

Hybride
CAD 70 000 - 110 000
Shared Services - Systems Administrator P2
Shared Services - Systems Administrator P2

OSI Maritime Systems Ltd. • Burnaby

Sur place
CAD 65 000 - 90 000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

Sur place
CAD 100 000 - 130 000
Systems Administrator
Systems Administrator

Zeektek • Stockton

Sur place
CAD 60 000 - 80 000
Senior Systems Administrator
Senior Systems Administrator

KuriosIT • Montreal (administrative region)

Sur place
CAD 65 000 - 90 000
Group insurance
Group RRSP
Flexible work policy
Systems Administrator
Systems Administrator

Kongsberg • Ottawa

Sur place
CAD 63 000 - 107 000
Accommodations available