Infrastructure Operations Engineer

Greenhouse Software, Inc.

Quebec

Hybrid

CAD 90,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Salary and discretionary bonus
Flexible remote/hybrid work
Ownership & autonomy
Career growth opportunities
Collaborative culture
International team environment
Impactful projects

Job summary

NexGen Cloud is expanding its OpenStack and Kubernetes footprint to support GPU workloads at scale. This role owns design, deployment, and operations of core infrastructure, with focus on performance, reliability, and security.

You will automate provisioning, drive incident response, and collaborate across Platform, DevOps, AI, Product, and Support teams. The position offers remote or hybrid work in a fast-growing, internationally minded company and opportunities to ship impactful infrastructure

Qualifications

  • Extensive Linux systems administration experience with depth.
  • Hands-on data centre hardware experience and on-site work.
  • Experience with OpenStack and Kubernetes at scale.
  • Familiarity with networking, storage, and security controls.
  • Willingness to travel to Quebec sites as required.

Responsibilities

  • Design, deploy, and operate OpenStack and Kubernetes environments for GPU workloads.
  • Automate provisioning and deployment using infrastructure-as-code and GitOps.
  • Improve workload scheduling and monitor platform stability with logging and alerts.
  • Lead incident response and drive reliability improvements.
  • Maintain RBAC, network policies, and tenant isolation across infra.
  • Collaborate with Platform, DevOps, AI, Product, and Support teams.

Skills

Linux admin
OpenStack
Kubernetes
Networking
GitOps
Incident response
RBAC & security
Hardware familiarity

Tools

OpenStack
Kubernetes
NVIDIA GPUs

Job description

NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature.

We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like.

THE ROLE: Infrastructure Operations Engineer

This role exists because our platform is scaling quickly — and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience.

This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately — you'll enjoy this.

WHAT YOU'LL BE DOING:

Rather than a long checklist, here's what success in this role looks like:

  • Own the design, deployment, and operation of OpenStack and Kubernetes environments — ensuring platform performance, scalability, and resilience for GPU workloads
  • Build and improve infrastructure using infrastructure-as-code and GitOps practices, driving automation across provisioning, deployment, and operational workflows
  • Optimise GPU workload scheduling using Kubernetes and NVIDIA tooling, and implement monitoring, logging, and alerting to ensure platform stability
  • Lead incident response and drive continuous improvement of reliability across the platform
  • Maintain strong security controls across infrastructure and container layers — RBAC, network policies, and tenant isolation
  • Work closely with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with customer and platform requirements
ABOUT YOU:

We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following:

  • Extensive hands‑on Linux systems administration skills and knowledge — genuine depth, not surface-level familiarity
  • Strong, proven experience building servers and racks — you've physically assembled, cabled, and commissioned hardware, not just specified or overseen it
  • Direct hands‑on experience physically working in data centres — you've personally stacked and racked hardware on‑site
  • A willingness and ability to travel to Quebec sites as required
  • A solid understanding of networking and storage systems

Nice to Have

  • Experience installing, racking, and configuring GPU hardware specifically, ideally including NVIDIA platforms
  • Production experience running OpenStack and/or Kubernetes at scale
  • Experience with infrastructure automation, CI/CD, and Git‑based workflows
  • Broader exposure to HPC or large‑scale compute environments
  • Contributions to open‑source projects
WHAT WE OFFER:
  • Competitive salary and annual discretionary bonus scheme
  • 25 days of holiday, plus public holidays
  • Flexible working arrangements (remote or hybrid, depending on role and location)
  • Real ownership and autonomy, with the trust to take initiative and experiment
  • The opportunity to make a visible, meaningful impact as we scale
  • Clear career progression and growth opportunities in a fast‑growing company
  • A collaborative, international culture built on trust, transparency, and ownership
  • The chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission‑driven colleagues
MORE INFORMATION

Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our journey, newest releases and hear exciting news in the neocloud space.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure Operations Engineer
Infrastructure Operations Engineer

NexGen Cloud • Quebec

On-site
CAD 100,000 - 140,000
Competitive salary
Wellbeing benefits
25 days of holiday
+1
Infrastructure Operations Engineer - GPU Cloud Platform
Infrastructure Operations Engineer - GPU Cloud Platform

Greenhouse Software, Inc. • Quebec

Hybrid
CAD 90,000 - 130,000
Salary and discretionary bonus
Flexible remote/hybrid work
Ownership & autonomy
+4
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.ai • Canada

On-site
CAD 120,000 - 180,000
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud

NVIDIA • Toronto

On-site
CAD 245,000 - 295,000
Equity
Benefits
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Remote Infrastructure/GPU Cluster/Platform Operations Lead
Remote Infrastructure/GPU Cluster/Platform Operations Lead

Bilinguallink • Lower Sackville

Remote
CAD 120,000 - 160,000
Remote Infrastructure/GPU Cluster/Platform Operations Lead
Remote Infrastructure/GPU Cluster/Platform Operations Lead

Bilinguallink • Brantford

Remote
CAD 140,000 - 190,000
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud

NVIDIA • Toronto

On-site
CAD 170,000 - 275,000
Equity
Benefits
GPU Cloud Platform Engineer
GPU Cloud Platform Engineer

Yotta Labs • Canada

On-site
CAD 90,000 - 120,000
Flexible work environment
Cutting-edge technology projects
Collaborative team culture
Cloud Platform Engineer
Cloud Platform Engineer

Swoon • Toronto

Hybrid
CAD 69,000 - 96,000