Platform Engineer (Remote)

ThinkOn Inc.

Toronto

Hybrid

CAD 110,000 - 115,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote-first culture
Competitive compensation
Flexible time-off
Health and dental benefits, GRSP & 401

Job summary

ThinkOn Inc. is hiring a Platform Engineer to build and maintain monitoring, observability, and infrastructure automation for our cloud platform. You will work with Zabbix, Prometheus/Grafana, Opsgenie, and IaC tooling in a VMware/Kubernetes environment.

We are a remote-first organization welcoming candidates from Canada. The role focuses on reliability, incident management, and collaborating across infra, network, security, and DevOps teams to support scalable platform operations.

Qualifications

  • Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience).
  • Eligible to obtain Secret Level Clearance within first 3 months; must be Canadian citizen with 10+ years background history.
  • Certifications such as Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations are a plus but not required.
  • Hands-on experience with Zabbix (or Nagios/Icinga/Checkmk).
  • Working knowledge of Prometheus and Grafana— writing exporters, dashboards, PromQL.
  • Experience with alert management and on-call tooling (Opsgenie, PagerDuty, or similar).
  • Comfort with Linux systems administration in a Linux-heavy environment.
  • Proficiency in scripting and automation— Bash and Python at minimum.
  • Experience with at least one IaC tool (Ansible, Terraform).
  • Familiarity with CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, or similar).
  • Basic database administration (PostgreSQL or MySQL) for monitoring backends.
  • Strong diagnostic and troubleshooting skills; clear written and verbal communication.
  • Attention to detail; ability to work independently in a remote environment.

Responsibilities

  • Deploy, configure, and maintain Zabbix for system and network monitoring across the platform.
  • Build and maintain Prometheus exporters and Grafana dashboards for capacity planning and visibility.
  • Configure and manage Opsgenie for alert routing, escalation, and on-call schedules.
  • Analyze alert noise, tune thresholds, and reduce false positives.
  • Integrate monitoring with ticketing, communication, and incident tools.
  • Maintain and optimize monitoring infrastructure incl. DB tuning and HA.
  • Write and maintain Ansible playbooks for deploying monitoring infra.
  • Use Terraform for provisioning where applicable.
  • Build and maintain GitLab CI/CD pipelines for monitoring and infra code.
  • Follow infrastructure-as-code practices with version control and peer review.
  • Act as first responder for monitoring incidents and alerts.
  • Investigate performance issues, outages, or anomalies from monitoring.
  • Escalate to network/security/infrastructure teams when needed.
  • Document incidents, root causes, and resolutions; contribute to post-incident reviews.
  • Provide technical support to internal teams on monitoring tools and dashboards.

Skills

Zabbix
Prometheus
Grafana
Opsgenie
Linux administration
Bash
Python
Ansible
Terraform
GitLab CI/CD
Kubernetes
VMware vCenter
Google Cloud/AWS/Linux tools

Education

Diploma or degree in Computer Science, IT, or related field

Tools

Ansible
Terraform
GitLab CI/CD

Job description

Salary Range: $110,000.00 To $115,000.00 Annually

As aPlatform Engineerat ThinkOn, you'll build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable. Your day-to-day will span Zabbix and Prometheus/Grafana for monitoring and dashboards, Opsgenie for alert routing and on-call management, and infrastructure-as-code tooling (Ansible, Terraform, GitLab CI/CD) to deploy and manage it all. You'll work within a VMware Cloud Foundation (VCF) and Kubernetes environment, collaborating with infrastructure, network, security, and DevOps teams.

ThinkOn is a remote-first organization. At this time, we are welcoming candidates from Canada for this position.

Please note that this listing is for a current vacancy at ThinkOn.

We are looking for qualified candidates who are eager to contribute and grow with us.

You Will:
Monitoring & Observability
  • Deploy, configure, and maintain Zabbix for system and network monitoring across the platform.
  • Build and maintain Prometheus exporters and Grafana dashboards for capacity planning, performance metrics, and operational visibility.
  • Configure and manage Ops genie for alert routing, escalation policies, and on-call schedules.
  • Analyze alert noise, tune thresholds, and reduce false positives to keep alerting actionable.
  • Integrate monitoring systems with ticketing, communication, and incident management tools.
  • Maintain and optimize monitoring infrastructure — database tuning, storage management, high availability.
Infrastructure Automation & CI/CD
  • Write and maintain Ansible playbooks for deploying and configuring monitoring infrastructure.
  • Use Terraform for provisioning infrastructure resources where applicable.
  • Build and maintain GitLab CI/CD pipelines for automated testing, linting, and deployment of monitoring and infrastructure code.
  • Follow infrastructure-as-code practices — version-controlled, peer-reviewed, reproducible.
  • Act as a first responder for monitoring-related incidents and alerts.
  • Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems.
  • Escalate to the appropriate teams (network, security, infrastructure) when needed.
  • Document incidents, root causes, and resolutions. Contribute to post-incident reviews.
  • Provide technical support to internal teams on monitoring tools and dashboards.
Platform Operations
  • Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs.
  • Manage notification infrastructure (SMTP relay configuration, delivery troubleshooting).
  • Support compliance requirements (ISO 27001, SOC 2) by maintaining audit logging, access controls, and security configurations for monitoring systems.
You Have:
  • Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience).
  • Eligible to obtain Secret Level Clearance within your first 3 months of employment. This requires the successful candidate to be a Canadian Citizen and have 10+ years of verifiable police background history.
  • Relevant certifications are a plus but not required (e.g., Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations).
  • Hands-on experience with Zabbix(or comparable: Nagios, Icinga,Checkmk).
  • Working knowledge of Prometheus and Grafana— writing exporters, building dashboards, PromQL.
  • Experience with alert management and on-call tooling(Ops genie, PagerDuty, or similar).
  • Comfort with Linux systems administration (this is a Linux-heavy environment).
  • Proficiency in scripting and automation— Bash and Python at minimum.
  • Experience with at least one IaCtool (Ansible, Terraform).
  • Familiarity with CI/CD pipelines(GitLab CI, GitHub Actions, Jenkins, or similar).
  • Basic database administration (PostgreSQL or MySQL) for monitoring tool backends.
  • Strong diagnostic and troubleshooting skills — you can work through a problem methodically.
  • Clear written and verbal communication —you'll document your work and explain technical issues to varied audiences.
  • Attention to detail — monitoring generates a lot of data, and you need to separate signal from noise.
  • Comfort working independently in a remote environment while collaborating across teams.
  • Ability to stay composed during incidents and work under time pressure.
Nice-to-Have
  • Experience with Kubernetes operations and troubleshooting.
  • Familiarity with VMware vSphere / VCF environments.
  • Exposure to log aggregation tools (ELK/OpenSearch, Loki, Graylog).
  • Knowledge of email security standards (SPF, DKIM, DMARC).
  • Experience cloud monitoring in multi-tenant or service-provider environments.
  • Understanding of Canadian data sovereignty requirements or public-sector compliance (PIPEDA, ITSG-33/PBMM).
  • Familiarity with ITSM platforms like Service Desk Plus by Managed Engine, ZenDesk , ServiceNow, Jira Service Management, or similar.
Benefits and Perks:
  • A remote-first culture
  • Competitive compensation package
  • Flexible time-off for vacation, plus illness & personal days
  • Comprehensive health and dental benefits, including GRSP & 401k Matching Program
About ThinkOn Inc. | Where Data Thrives:

Founded in 2013, ThinkOn is a managed infrastructure services provider (MISP) with a global data center footprint, focused on empowering partners to do whatever they need to do with their data. ThinkOn’s team of data-obsessed experts protect clients’ data like it’s their own, making it more resilient, secure, actionable, and searchable. ThinkOn is Channel First and works with a global network of value-add resellers and managed service providers to provide creative, turnkey Infrastructure-as-a-Service (IaaS), Disaster Recovery-as-a-Service (DRaaS), and Backup-as-a-Service (BaaS) solutions and data management services that are fast, flexible, scalable, highly secure, and cost-effective with predictable pricing and no hidden fees. ThinkOn has data centers located across North America, the United Kingdom, Australia, and the Caribbean.

To learn more, visitwww.ThinkOn.com

Accessibility Accommodations:

ThinkOn is committed to a workforce that is reflective of diverse populations. We welcome applications from qualified individuals from all backgrounds. In accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and accessibility standards across Canada, ThinkOn provides accommodations to job applicants with disabilities throughout the recruitment process. If you require accommodations, please let us know and we will work with you to meet your needs. We are committed to a selection process and work environment that is inclusive, equitable, accessible, and adheres to our corporate values.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Platform Engineer
Cloud Platform Engineer

ThinkOn Inc. • Toronto

Hybrid
CAD 110,000 - 125,000
Remote-first culture
Competitive pay
Flexible time-off
+4
IT Network Technician
IT Network Technician

ThinkOn Inc. • Toronto

Hybrid
CAD 55,000 - 65,000
Remote-first culture
Competitive compensation package
Flexible time-off
+3
Jr. Data Centre Technician
Jr. Data Centre Technician

ThinkOn Inc. • Toronto

Hybrid
CAD 50,000 - 55,000
Competitive compensation
Flexible time-off
Health & dental benefits
+3
Jr. Data Centre Technician
Jr. Data Centre Technician

Socket.dev • Toronto

On-site
CAD 42,000 - 62,000
Competitive compensation package
Flexible time-off for vacation; 5 days
Comprehensive health and dental benefit
+3
Service Desk Analyst
Service Desk Analyst

Thinkon Inc • North Bay

Hybrid
CAD 55,000 - 59,000
Service Desk Analyst
Service Desk Analyst

ThinkOn Inc. • North Bay

Hybrid
CAD 55,000 - 59,000
Service Desk Analyst
Service Desk Analyst

Socket.dev • North Bay

Hybrid
CAD 42,000 - 64,000
Linux Systems Administrator
Linux Systems Administrator

STACK IT Recruitment • Vaughan

On-site
CAD 90,000 - 110,000
Competitive vacation and personal days
Health Spending Account
Dental Coverage
+1
Systems Engineer
Systems Engineer

STACK IT Recruitment • Toronto

On-site
CAD 80,000 - 90,000
Benefits Package
RRSP Matching
Paid Time Off
Manager, Systems Engineering
Manager, Systems Engineering

ServiceNow • Toronto

On-site
CAD 126,000 - 220,000
Health plans
RRSP with company match
ESPP
+3