Observability Engineer

Apptoza Inc.

Toronto

On-site

CAD 120,000 - 170,000

Full time

5 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Apptoza Inc. seeks an experienced Observability Engineer to own and advance the observability stack across 50+ production Kubernetes clusters in a Toronto-based setup.

You will implement metrics, logging, tracing, and alerting pipelines with GitOps and IaC, while driving AI-enhanced monitoring capabilities. You will collaborate with SRE and platform teams to ensure reliability, performance, and cost efficiency across multi-cloud environments, employing Prometheus, Grafana, and modern storage

Qualifications

  • 4–6 years in observability, monitoring, SRE, or platform engineering roles.
  • 3+ years hands-on production experience with Prometheus and Grafana.
  • Experience with Kubernetes and containerized environments.
  • Strong PromQL, log aggregation, and tracing skills.
  • Experience deploying and managing observability stacks at scale (1000+ nodes).

Responsibilities

  • Own the observability stack across 50+ production Kubernetes clusters.
  • Design, deploy, and maintain metrics, logging, tracing, and alerting capabilities.
  • Implement GitOps-based deployments and IaC.
  • Develop dashboards with SLO/SLI tracking and executive metrics.
  • Collaborate with SRE, DevOps and platform teams.
  • Enable AI-powered anomaly detection and predictive alerting.

Skills

SRE practices
Incident management
Dashboard design

Tools

Prometheus
Grafana
Thanos
Loki
Promtail
Fluentd
OpenTelemetry
Jaeger
GitOps

Job description

Role: Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)
Location: Toronto, ON – 4 days onsite
Contract Role
Job Description
ABOUT THE ROLE

We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.

OBSERVABILITY STACK OWNERSHIP
  • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents
  • Manage observability deployments using GitOps principles and infrastructure as code
  • Implement long-term metrics storage solutions with cloud object storage
  • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments
  • Configure distributed observability architecture spanning multiple data centers and cloud providers
METRICS & MONITORING
  • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications
  • Create ServiceMonitors, PodMonitors for automated metrics collection
  • Develop rules for intelligent alerting with minimal false positives
  • Configure multi-cluster metrics federation and aggregation
  • Optimize metrics cardinality, storage efficiency, and query performance
  • Implement recording rules for pre-aggregated metrics and SLI calculations
DASHBOARDS & VISUALIZATION
  • Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics
  • Create reusable dashboard templates and libraries for development teams
  • Implement dashboard-as-code
  • Configure multiple datasources
  • Design executive dashboards with SLO/SLI tracking and business KPIs
  • Implement role-based access control and multi-tenancy in Grafana
LOGGING INFRASTRUCTURE
  • Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or Similar)
  • Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.)
  • Design log retention policies balancing cost, compliance, and operational needs
  • Create LogQL/Lucene queries and log-based alerts
  • Implement log correlation with metrics and traces for unified troubleshooting
  • Build log aggregation pipelines with parsing, filtering, and enrichment
ALERTING & INCIDENT MANAGEMENT
  • Configure intelligent alerting with Alertmanager or equivalent platforms
  • Design alert rules with appropriate severity levels, thresholds, and SLOs
  • Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks
  • Create automated runbooks and remediation workflows
  • Develop alert inhibition, silencing, and grouping strategies
  • Tune alerting to achieve signal-to-noise ratio improvements
  • Integrate with incident management and on-call rotation systems
DISTRIBUTED TRACING
  • Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo
  • Instrument applications for trace collection and correlation
  • Configure trace sampling strategies for cost and performance optimization
  • Build trace-based dashboards for latency analysis and dependency mapping
  • Integrate tracing with metrics and logs for comprehensive observability
AI/ML FOR OBSERVABILITY
  • Implement AI-powered anomaly detection for metrics and logs
  • Build predictive alerting using machine learning models to forecast issues before they occur
  • Develop intelligent alert correlation and root cause analysis systems
  • Integrate LLM-based tools for log analysis and troubleshooting assistance
  • Implement AIOps capabilities for automated incident triage and resolution
  • Use AI to optimize alert thresholds and reduce false positives
  • Build natural language query interfaces for observability data
  • Implement Model Context Protocol (MCP) for AI agent integration with observability platforms
AUTOMATION & PLATFORM INTEGRATION
  • Automate observability deployment using GitOps workflows (FluxCD, ArgoCD)
  • Integrate with CI/CD pipelines for automated testing and validation
  • Build self-service portals for teams to create dashboards and alerts
  • Develop APIs and CLIs for observability automation
  • Integrate with secrets management solutions (Vault, AWS Secrets Manager)
  • Configure LDAP/AD/SSO authentication for observability platforms
  • Automate compliance reporting and audit logging
ENABLEMENT & COLLABORATION
  • Onboard application teams to observability platforms
  • Provide guidance on instrumentation best practices
  • Create documentation, training materials, and self-service guides
  • Conduct workshops
  • Support development teams during incidents with observability insights
  • Collaborate with SRE, DevOps, and platform engineering teams
PERFORMANCE & COST OPTIMIZATION
  • Monitor and optimize observability stack resource consumption
  • Implement autoscaling for stateless observability components
  • Tune data retention, compaction, and downsampling strategies
  • Conduct capacity planning for metrics and log storage growth
  • Optimize query performance and dashboard response times
  • Implement cost allocation and chargeback for multi-tenant environments
REQUIRED QUALIFICATIONS
EXPERIENCE
  • 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles
  • 3+ years hands-on production experience with Prometheus and Grafana
  • 2+ years working with Kubernetes and containerized environments
  • Strong expertise with PromQL for metrics querying and alerting
  • Experience deploying and managing observability stacks at scale (1000+ nodes)
  • Proven track record of reducing MTTR through effective observability
CORE TECHNICAL SKILLS
OBSERVABILITY PLATFORMS:
  • Prometheus (including Prometheus Operator)
  • Grafana (dashboards, alerting, plugins)
  • Thanos, Cortex, or Mimir for long-term storage
  • Alertmanager or equivalent alerting platforms
LOGGING:
  • Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs
  • Log collection agents (Promtail, Fluentd, FluentBit, Vector)
  • LogQL, Lucene, or equivalent query languages
TRACING:
  • OpenTelemetry (OTEL Collector, instrumentation)
  • Jaeger, Zipkin, Tempo, or AWS X-Ray
  • Trace sampling and correlation strategies
KUBERNETES & CONTAINERS:
  • Kubernetes architecture and operations (1.24+)
  • Custom Resource Definitions (CRDs) and Operators
  • ServiceMonitor, PodMonitor, PrometheusRule resources
  • Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor)
GITOPS & AUTOMATION:
  • Basic SQL for data analysis
CLOUD & INFRASTRUCTURE:
  • AWS, Azure, or GCP cloud platforms
  • Multi-cloud and hybrid architectures
  • vSphere or on-premises virtualization (nice to have)
AI/ML & EMERGING TECHNOLOGIES
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)
Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)

Apptoza Inc. • Toronto

Hybrid
CAD 120,000 - 160,000
Senior Observability Engineer
Senior Observability Engineer

Apptoza Inc. • Toronto

Hybrid
CAD 130,000 - 180,000
Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

Hybrid
CAD 120,000 - 160,000
GCP Observability Engineer
GCP Observability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 85,000 - 115,000
Senior Site Reliability Engineer (Observability & Analytics, Platform Infra)
Senior Site Reliability Engineer (Observability & Analytics, Platform Infra)

Elastic • Ottawa

Hybrid
CAD 120,000 - 170,000
Health coverage for you and family
Flexible location and schedule
Generous vacation days
+3
Platform & SRE Engineer
Platform & SRE Engineer

TechDoQuest • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Dynatrace Observability Platform Engineer
Dynatrace Observability Platform Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Mississauga

On-site
CAD 90,000 - 140,000
Site Reliability Engineer (SRE) – Observability
Site Reliability Engineer (SRE) – Observability

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

On-site
CAD 75,000 - 95,000
Founding Engineer (Agentic Platform)
Founding Engineer (Agentic Platform)

Katalyze AI, Inc. • Toronto

On-site
CAD 130,000 - 200,000