Remote Infrastructure/GPU Cluster/Platform Operations Lead

Bilinguallink

Lower Sackville

Remote

CAD 120,000 - 160,000

Full time

44 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ELEKS is seeking an Infrastructure/GPU Cluster/Platform Operations Lead to oversee GPU-driven infrastructure in Canada. The role focuses on Kubernetes-based AI environments, cloud integration, and reliability practices to support mission-critical AI deployments.

The candidate will lead design, capacity planning, and security strategies while collaborating with AI engineering teams and driving platform evolution in a fast-paced environment.

Qualifications

  • 8+ years of Infrastructure Engineering or Platform Operations experience.
  • Experience managing GPU clusters for AI workloads.
  • Strong Kubernetes administration skills.
  • Experience with NVIDIA GPU technologies and CUDA ecosystem.
  • Experience with cloud infrastructure (Azure, AWS or GCP).
  • Knowledge of storage, networking, and high-performance computing environments.
  • Experience implementing Infrastructure as Code (Terraform or similar).
  • Strong operational leadership skills.
  • Experience supporting AI platform infrastructure.
  • Upper-Intermediate or higher level of English.

Responsibilities

  • Lead GPU infrastructure design and operations.
  • Manage Kubernetes-based AI platform environments.
  • Optimize infrastructure for AI training and inference workloads.
  • Define operational standards and reliability practices.
  • Collaborate with AI engineering teams.
  • Implement monitoring, security, and disaster recovery strategies.
  • Lead infrastructure capacity planning.
  • Support technical roadmap and infrastructure evolution.

Skills

Infra engineering
GPU clusters
Kubernetes
NVIDIA CUDA
Cloud platforms (Azure/AWS/GCP)
Storage & networking
Infrastructure as Code (Terraform)
Leadership
AI platform infra support
English (Upper-Intermediate+)

Tools

Terraform
CUDA toolkit

Job description

ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada.
Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered.

ABOUT CLIENT

Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making.

The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data.

REQUIREMENTS
  • 8+ years of Infrastructure Engineering or Platform Operations experience
  • Experience managing GPU clusters for AI workloads
  • Strong Kubernetes administration skills
  • Experience with NVIDIA GPU technologies and CUDA ecosystem
  • Experience with cloud infrastructure (Azure, AWS or GCP)
  • Knowledge of storage, networking, and high-performance computing environments
  • Experience implementing Infrastructure as Code (Terraform or similar)
  • Strong operational leadership skills
  • Experience supporting AI platform infrastructure
  • Upper-Intermediate or higher level of English
RESPONSIBILITIES
  • Lead GPU infrastructure design and operations
  • Manage Kubernetes-based AI platform environments
  • Optimize infrastructure for AI training and inference workloads
  • Define operational standards and reliability practices
  • Collaborate with AI engineering teams
  • Implement monitoring, security, and disaster recovery strategies
  • Lead infrastructure capacity planning
  • Support technical roadmap and infrastructure evolution

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Infrastructure/GPU Cluster/Platform Operations Lead
Remote Infrastructure/GPU Cluster/Platform Operations Lead

Bilinguallink • Brantford

Remote
CAD 140,000 - 190,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Senior Engineer-Cloud AI Infrastructure
Senior Engineer-Cloud AI Infrastructure

Huawei Technologies Co. • Markham

Hybrid
CAD 177,000 - 314,000
Infrastructure Operations Engineer
Infrastructure Operations Engineer

Greenhouse Software, Inc. • Quebec

Hybrid
CAD 90,000 - 130,000
Salary and discretionary bonus
Flexible remote/hybrid work
Ownership & autonomy
+4
AI/ML Infrastructure Engineer
AI/ML Infrastructure Engineer

BULL-IT SOLUTIONS LTD • Montreal

On-site
CAD 100,000 - 130,000
Senior IT Application Specialist, DevOps & AI Operations - SaaS Client
Senior IT Application Specialist, DevOps & AI Operations - SaaS Client

S.I. Systems Ltd. • Edmonton, Calgary, Vancouver, Victoria, Ottawa

Hybrid
CAD 110,000 - 140,000
Infrastructure Operations Engineer
Infrastructure Operations Engineer

NexGen Cloud • Quebec

On-site
CAD 100,000 - 140,000
Competitive salary
Wellbeing benefits
25 days of holiday
+1
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud

NVIDIA • Toronto

On-site
CAD 245,000 - 295,000
Equity
Benefits
Research Engineer - AI Workload & Systems
Research Engineer - AI Workload & Systems

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 178,070 - 315,479
Data Center Operations Lead
Data Center Operations Lead

IREN • Prince George

On-site
CAD 70,000 - 90,000
Competitive hourly rate
RRSP matching program
Relocation assistance
+4