Observability Engineer – Kubernetes Platform

BullTechSystem

Halifax

On-site

CAD 90,000 - 120,000

Full time

46 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

BullTechSystem is seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team. You’ll own the complete observability stack across 50+ production Kubernetes clusters, delivering metrics, logging, tracing, and alerts to ensure reliability for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing

Qualifications

  • Experience designing and operating observability stacks at scale.
  • Proficiency with Prometheus, Grafana, Thanos and Loki.
  • Strong experience with GitOps and IaC for deployments.

Responsibilities

  • Design, deploy, and maintain enterprise-scale observability infrastructure across 50+ Kubernetes clusters.
  • Implement long-term metrics storage with cloud object storage.
  • Maintain observability components across dev, QA, UAT, prod, and DR environments.
  • Enable intelligent monitoring with AI/ML capabilities for predictive alerting.

Skills

Prometheus
Grafana
Thanos
Loki
Kubernetes
GitOps

Tools

Argo CD
Terraform
Cloud object storage

Job description

ABOUT THE ROLE

We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization.

You’ll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.

WHAT YOULL DO
  • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents
  • Manage observability deployments using GitOps principles and infrastructures code
  • Implement long-term metrics storage solutions with cloud object storage
  • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments
  • Configure distributed observability architecture spanning multiple datacenters and cloud providers
METRICS & MONITORING
  • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications
  • Create Service Monitors, Pod Monitors for automated metrics collection
  • Develop rules for intelligent alerting with minimal false positives
  • Configure multi-cluster metrics federation and aggregation
  • Optimize metrics cardinality, storage deficiency, and query performance.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

Apptoza Inc. • Toronto

On-site
CAD 120,000 - 170,000
Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000
Senior Observability Engineer for Multi-Cluster Kubernetes
Senior Observability Engineer for Multi-Cluster Kubernetes

BullTechSystem • Halifax

On-site
CAD 90,000 - 120,000
Integration Engineer
Integration Engineer

Build Nova Scotia • Halifax

On-site
CAD 85,000 - 115,000
GCP Observability Engineer
GCP Observability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 85,000 - 115,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

On-site
CAD 120,000 - 160,000
Platform & SRE Engineer
Platform & SRE Engineer

TechDoQuest • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Cloud Observability Engineer
Cloud Observability Engineer

TMC • Montreal (administrative region)

On-site
CAD 80,000 - 100,000
Permanent employment contract
Profit sharing
One-on-one coaching and training
+2
Platform Engineer
Platform Engineer

Hunter Bond • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Dynatrace Observability Platform Engineer
Dynatrace Observability Platform Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Mississauga

On-site
CAD 90,000 - 140,000