Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)

Apptoza Inc.

Toronto

Hybrid

CAD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Apptoza Inc. is seeking an experienced Observability Engineer to own the enterprise observability stack across 50+ Kubernetes clusters.

You will design, deploy, and maintain metrics, logs, tracing, and alerting with Prometheus, Grafana, Loki, and related tools, while applying GitOps and IaC to enable scalable, self-healing infrastructure. You will drive intelligent monitoring using AI/ML capabilities and collaborate with SRE, DevOps, and platform teams to optimize performance, reliability, and

Qualifications

  • 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles.
  • 3+ years hands-on production experience with Prometheus and Grafana.
  • 2+ years working with Kubernetes and containerized environments.
  • Strong expertise with PromQL for metrics querying and alerting.
  • Experience deploying and managing observability stacks at scale (1000+ nodes).
  • Proven track record of reducing MTTR through effective observability.

Responsibilities

  • Design, deploy, and maintain enterprise-scale observability infrastructure across clusters and environments.
  • Manage observability deployments using GitOps principles and infrastructure as code
  • Create dashboards and alerts, and enable multi-datasource visualization in Grafana
  • Configure logging and tracing pipelines to integrate with metrics
  • Collaborate with SRE, DevOps, and platform teams to optimize reliability and performance

Skills

Observability
PromQL
Kubernetes
SRE practices
AI/ML familiarity

Tools

Prometheus
Grafana
Thanos
Loki
OpenTelemetry
Jaeger
Tempo
Elasticsearch

Job description

Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)

Toronto, On

Hybrid - 4 days

ABOUT THE ROLE

We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.

WHAT YOU'LL DO
OBSERVABILITY STACK OWNERSHIP
  • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents
  • Manage observability deployments using GitOps principles and infrastructure as code
  • Implement long-term metrics storage solutions with cloud object storage
  • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments
  • Configure distributed observability architecture spanning multiple data centers and cloud providers
METRICS & MONITORING
  • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications
  • Create ServiceMonitors, PodMonitors for automated metrics collection
  • Develop rules for intelligent alerting with minimal false positives
  • Configure multi-cluster metrics federation and aggregation
  • Optimize metrics cardinality, storage e_iciency, and query performance
  • Implement recording rules for pre-aggregated metrics and SLI calculations
DASHBOARDS & VISUALIZATION
  • Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics
  • Create reusable dashboard templates and libraries for development teams
  • Implement dashboard-as-code
  • Configure multiple datasources
  • Design executive dashboards with SLO/SLI tracking and business KPIs
  • Implement role-based access control and multi-tenancy in Grafana
LOGGING INFRASTRUCTURE
  • Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or similar)
  • Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.)
  • Design log retention policies balancing cost, compliance, and operational needs
  • Create LogQL/Lucene queries and log-based alerts
  • Implement log correlation with metrics and traces for unified troubleshooting
  • Build log aggregation pipelines with parsing, filtering, and enrichment
ALERTING & INCIDENT MANAGEMENT
  • Configure intelligent alerting with Alertmanager or equivalent platforms
  • Design alert rules with appropriate severity levels, thresholds, and SLOs
  • Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks
  • Create automated runbooks and remediation workflows
  • Develop alert inhibition, silencing, and grouping strategies
  • Tune alerting to achieve signal-to-noise ratio improvements
  • Integrate with incident management and on-call rotation systems
DISTRIBUTED TRACING
  • Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo
  • Instrument applications for trace collection and correlation
  • Configure trace sampling strategies for cost and performance optimization
  • Build trace-based dashboards for latency analysis and dependency mapping
  • Integrate tracing with metrics and logs for comprehensive observability
AI/ML FOR OBSERVABILITY
  • Implement AI-powered anomaly detection for metrics and logs
  • Build predictive alerting using machine learning models to forecast issues before they occur
  • Develop intelligent alert correlation and root cause analysis systems
  • Integrate LLM-based tools for log analysis and troubleshooting assistance
  • Implement AIOps capabilities for automated incident triage and resolution
  • Use AI to optimize alert thresholds and reduce false positives
  • Build natural language query interfaces for observability data
  • Implement Model Context Protocol (MCP) for AI agent integration with observability platforms
AUTOMATION & PLATFORM INTEGRATION
  • Automate observability deployment using GitOps workflows (FluxCD, ArgoCD)
  • Integrate with CI/CD pipelines for automated testing and validation
  • Build self-service portals for teams to create dashboards and alerts
  • Develop APIs and CLIs for observability automation
  • Integrate with secrets management solutions (Vault, AWS Secrets Manager)
  • Configure LDAP/AD/SSO authentication for observability platforms
  • Automate compliance reporting and audit logging
ENABLEMENT & COLLABORATION
  • Onboard application teams to observability platforms
  • Provide guidance on instrumentation best practices
  • Create documentation, training materials, and self-service guides
  • Conduct workshops
  • Support development teams during incidents with observability insights
  • Collaborate with SRE, DevOps, and platform engineering teams
PERFORMANCE & COST OPTIMIZATION
  • Monitor and optimize observability stack resource consumption
  • Implement autoscaling for stateless observability components
  • Tune data retention, compaction, and downsampling strategies
  • Conduct capacity planning for metrics and log storage growth
  • Optimize query performance and dashboard response times
  • Implement cost allocation and chargeback for multi-tenant environments
REQUIRED QUALIFICATIONS
EXPERIENCE
  • 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles
  • 3+ years hands-on production experience with Prometheus and Grafana
  • 2+ years working with Kubernetes and containerized environments
  • Strong expertise with PromQL for metrics querying and alerting
  • Experience deploying and managing observability stacks at scale (1000+ nodes)
  • Proven track record of reducing MTTR through e_ective observability
CORE TECHNICAL SKILLS
OBSERVABILITY PLATFORMS:
  • Prometheus (including Prometheus Operator)
  • Grafana (dashboards, alerting, plugins)
  • Thanos, Cortex, or Mimir for long-term storage
  • Alertmanager or equivalent alerting platforms
LOGGING:
  • Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs
  • Log collection agents (Promtail, Fluentd, FluentBit, Vector)
  • LogQL, Lucene, or equivalent query languages
TRACING:
  • OpenTelemetry (OTEL Collector, instrumentation)
  • Jaeger, Zipkin, Tempo, or AWS X-Ray
  • Trace sampling and correlation strategies
KUBERNETES & CONTAINERS:
  • Kubernetes architecture and operations (1.24+)
  • Custom Resource Definitions (CRDs) and Operators
  • ServiceMonitor, PodMonitor, PrometheusRule resources
  • Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor)
GITOPS & AUTOMATION:
  • Basic SQL for data analysis
CLOUD & INFRASTRUCTURE:
  • AWS, Azure, or GCP cloud platforms
  • Multi-cloud and hybrid architectures
  • vSphere or on-premises virtualization (nice to have)
AI/ML & EMERGING TECHNOLOGIES
  • Experience with AI/ML frameworks for observability (Prophet, TensorFlow, PyTorch)
  • Anomaly detection algorithms and time-series forecasting
  • LLM integration for log analysis and troubleshooting (GPT, Claude, etc.)
  • Model Context Protocol (MCP) for AI agent integration
  • AIOps platforms (Moogsoft, BigPanda, Datadog Watchdog, etc.)
  • Natural language processing for log parsing and analysis
  • Familiarity with vector databases for semantic search (Pinecone, Weaviate)
  • Experience with AI-powered root cause analysis tools
  • Knowledge of prompt engineering for observability use cases
PREFERRED QUALIFICATIONS
  • Kubernetes certifications (CKA, CKAD, or CKS)
  • Grafana certification or equivalent training
  • Knowledge of service mesh observability (
  • Experience with SRE practices, SLIs, SLOs, and error budgets
  • Background in financial services or regulated industries
  • Familiarity with compliance requirements (SOX, PCI-DSS, etc.)
  • Contributions to open-source observability projects
  • Experience with APM tools
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Observability Engineer
Senior Observability Engineer

Apptoza Inc. • Toronto

Hybrid
CAD 130,000 - 180,000
Observability Engineer
Observability Engineer

Apptoza Inc. • Toronto

On-site
CAD 120,000 - 170,000
Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000
Senior Observability Engineer
Senior Observability Engineer

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Montreal (administrative region)

Hybrid
CAD 120,000 - 160,000
GCP Observability Engineer
GCP Observability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 85,000 - 115,000
Platform & SRE Engineer
Platform & SRE Engineer

TechDoQuest • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech • Toronto

On-site
CAD 120,000 - 180,000
Senior Infrastructure SRE
Senior Infrastructure SRE

Jobtailor • Mississauga

On-site
CAD 110,000 - 160,000
Senior Site Reliability Engineer (Observability & Analytics, Platform Infra)
Senior Site Reliability Engineer (Observability & Analytics, Platform Infra)

Elastic • Ottawa

Hybrid
CAD 120,000 - 170,000
Health coverage for you and family
Flexible location and schedule
Generous vacation days
+3
Senior Full-Stack Software Engineer
Senior Full-Stack Software Engineer

Jobtailor • Toronto

On-site
CAD 130,000 - 185,000