System Administrator

MGIS Inc.

Canada

On-site

CAD 80,000 - 100,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

MGIS Inc. is seeking a System Administrator, Level 2, in Canada to manage High Performance Computing (HPC) clusters, assisting scientists with user support. This role covers troubleshooting and maintaining systems that include CPU/GPU nodes, job schedulers, and parallel storage.

The ideal candidate will have solid experience in Linux HPC clusters administration and must be familiar with scientific computing toolchains. Eligible candidates need to be able to maintain a Secret-level security clearance.

Qualifications

  • Solid experience administering Linux-based HPC clusters.
  • Hands-on experience with job schedulers.
  • Experience troubleshooting CUDA installations and GPU failures.
  • Familiarity with scientific computing toolchains.

Responsibilities

  • Maintain the HPC cluster and perform hardware management.
  • Troubleshoot incidents to ensure quick recovery.
  • Support application builds and runtime troubleshooting.
  • Document processes and submit progress reports.

Skills

Linux-based HPC administration
CUDA troubleshooting
Job schedulers (PBS Pro/Torque, SLURM, SGE)
Scientific computing toolchains
Configuration management tools

Tools

Git
Ansible
MS DevOps

Job description

MGIS is seeking a System Administrator, Level 2, to manage High Performance Computing (HPC) clusters and support the scientists who rely on them. This role blends HPC system administration with hands‑on user support — helping researchers install, run, and debug applications on HPC infrastructure so they can focus on their science instead of IT issues.

HPC environments in scope include clustered CPU/GPU systems with job schedulers and attached parallel storage (e.g., Lustre, GPFS).

What you'll be doing
  • Maintain the HPC cluster — hardware, image management, local networking, scheduler, and backups
  • Troubleshoot environment incidents to ensure a quick return to normal operations
  • Meet with scientists to evaluate their HPC support requirements
  • Develop task plans to meet researchers' needs, consulting the technical authority for approval
  • Support application builds, installs, and runtime troubleshooting (GNU, Intel, Fortran, Nvidia)
  • Support open‑source and commercial software, including Python/Anaconda installs, Bash scripting, build/make tools, EasyBuild, Spack, and MPI implementations (MPICH, OpenMPI, IntelMPI, HPMPI)
  • Assist with compilation and runtime of in‑house developed applications
  • General systems management: Linux OS patching schedules and reliability
  • Manage user accounts (creation, deletion) and environment modules
  • Manage configuration via Git, MS DevOps, and Ansible Playbooks
  • Manage RPM/DEB packages and troubleshoot ThinLinc
  • Troubleshoot jobs on schedulers (PBS Pro/Torque, SLURM, SGE)
  • Ensure reliable CUDA installs; troubleshoot GPU failures and CUDA software/driver issues
  • Provide hardware support — memory upgrades, storage arrays, power/network cabling, ILO
  • Document every process and task to support enterprise knowledge continuity
  • Submit weekly progress reports to the Technical Authority
Requirements
  • Solid experience administering Linux-based HPC clusters (CPU/GPU nodes, schedulers, parallel storage)
  • Hands‑on experience with job schedulers such as PBS Pro/Torque, SLURM, or SGE
  • Experience troubleshooting CUDA installations, GPU failures, and driver issues
  • Familiarity with scientific computing toolchains — compilers (GNU, Intel), MPI implementations, EasyBuild, and Spack
  • Experience supporting researchers or end‑users with application builds and runtime issues
  • Working knowledge of configuration management tools (Git, Ansible, MS DevOps)
  • Comfortable working independently and producing clear technical documentation
  • Eligible to obtain and maintain a Secret‑level security clearance
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Specialist
HPC Specialist

Arcadion • Ottawa

On-site
CAD 80,000 - 110,000
SLURM HPC Architect / Administrator
SLURM HPC Architect / Administrator

Arcadion • Ottawa

On-site
CAD 90,000 - 120,000
Competitive compensation
Flexible employment structure
Remote-first environment
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Shared Services - Systems Administrator P2
Shared Services - Systems Administrator P2

OSI Maritime Systems Ltd. • Burnaby

On-site
CAD 65,000 - 90,000
Shared Services - Systems Administrator P2
Shared Services - Systems Administrator P2

Osimaritime • Burnaby

On-site
CAD 65,000 - 95,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Insight Softmax Consulting, LLC. • Canada

Remote
CAD 120,000 - 180,000
Senior DevOps Engineer HPC / EDA / SLURM- #26-24242
Senior DevOps Engineer HPC / EDA / SLURM- #26-24242

JobDiva, Inc. • Vancouver

On-site
CAD 110,000 - 140,000
Systems Administrator
Systems Administrator

Zeektek • Stockton

On-site
CAD 60,000 - 80,000
HPC Specialist
HPC Specialist

DRW Holdings, LLC. • Montreal (administrative region)

On-site
CAD 100,000 - 130,000
Senior Systems Administrator
Senior Systems Administrator

KuriosIT • Montreal (administrative region)

On-site
CAD 65,000 - 90,000
Group insurance
Group RRSP
Flexible work policy