Observability Specialist

Jobgether

Canada

Remote

CAD 110,000 - 150,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work opportunity

Job summary

Jobgether in Canada is seeking an Observability Specialist to shape observability for large-scale financial technology systems serving millions of users across Africa. You will own and evolve an observability platform spanning backend services, APIs, databases, Kubernetes, cloud infrastructure, and on-premises environments.

You will partner with engineering and platform teams to define SLIs, improve alert quality, and build internal tools that enable instrumentation, incident analysis, and

Qualifications

  • 5+ years of experience in observability or SRE.
  • Experience with distributed tracing, metrics, logs, and incident response at scale.
  • Proficiency in Python and OpenTelemetry instrumentation.
  • Strong communication and collaboration skills.

Responsibilities

  • Own, operate, and evolve the observability platform across apps, APIs, databases, Kubernetes workloads, cloud infra and on-prem environments.
  • Improve visibility into production behavior with metrics, logs, traces, dashboards, and SLIs.
  • Define meaningful SLIs and improve alert quality while connecting alerts to user impact.
  • Build internal tools and self-service workflows for instrumentation, incident analysis, and performance tuning.
  • Identify reliability, latency, capacity, performance, and cost issues before users are affected.
  • Establish observability standards and documentation adopted across teams.
  • Manage Datadog, Honeycomb, Sentry, OpenTelemetry, Prometheus, Grafana, and related platforms.
  • Support design and adoption of SLOs and reliability practices.
  • Investigate production issues and performance regressions across systems.
  • Collaborate with product, database, infrastructure, security, and engineering teams.

Skills

Python
Observability concepts
Backend engineering
Team collaboration
Autonomy

Tools

Prometheus
Grafana
Datadog
OpenTelemetry
Jaeger
Tempo
Loki
Honeycomb
Sentry
Pyroscope
PostgreSQL
CockroachDB
Redis
GraphQL
Kubernetes

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Observability Specialist based in Canada.

This is an opportunity to shape observability for large-scale financial technology systems serving millions of users across Africa.
You will own and evolve an observability platform spanning backend services, APIs, databases, Kubernetes, cloud infrastructure, and on-premises environments.
Your work will help engineering teams detect regressions earlier, investigate incidents faster, and make smarter reliability and performance decisions.
You’ll join a newly established Performance & Observability team with significant room to define standards, tooling, and ways of working.
The role combines hands-on engineering, platform ownership, automation, performance analysis, and close collaboration with product and infrastructure teams.
You’ll have strong autonomy, opportunities to mentor engineers, and the chance to solve complex problems at meaningful scale.
The environment values pragmatic engineering, fast iteration, simplicity, and technology that delivers measurable impact.

Accountabilities:
  • Own, operate, and continuously evolve the observability platform across applications, APIs, databases, Kubernetes workloads, cloud infrastructure, and on-premises environments.
  • Improve visibility into production behavior through effective metrics, logs, traces, profiling, dashboards, alerts, and service-level indicators.
  • Partner with engineering and platform teams to define meaningful SLIs and improve alert quality, reducing noise while keeping alerts actionable and connected to user impact.
  • Build internal tools, libraries, automation, and self-service workflows that enable engineers to instrument services, investigate incidents, analyze performance, and understand system dependencies.
  • Identify reliability, latency, capacity, performance, and infrastructure cost issues before they become user-facing incidents.
  • Establish observability standards, naming conventions, documentation, and training materials that can be adopted consistently across engineering teams.
  • Manage and optimize observability technologies such as Datadog, Honeycomb, Sentry, Pyroscope, Prometheus, Grafana, and OpenTelemetry, balancing scalability with platform costs.
  • Support the design and implementation of SLOs and help product teams adopt effective reliability practices.
  • Investigate production issues and performance regressions, including profiling and analysis of application and infrastructure behavior.
  • Collaborate with product, database, infrastructure, security, and engineering teams to improve system reliability and operational maturity.
Requirements:
  • 5+ years of experience in observability, SRE, platform engineering, infrastructure engineering, backend engineering, or production systems engineering.
  • Strong understanding of metrics, logging, distributed tracing, profiling, alerting, dashboards, SLIs/SLOs, and incident response workflows at scale.
  • Experience building internal tools, automation, libraries, or platforms used by other engineers.
  • Proficiency in at least one backend programming language, preferably Python.
  • Hands-on experience with observability platforms such as Prometheus, Grafana, Datadog, OpenTelemetry, Jaeger, Tempo, Loki, Honeycomb, Sentry, Pyroscope, or similar technologies.
  • Experience with OpenTelemetry instrumentation and collector configuration at scale.
  • Familiarity with technologies such as PostgreSQL, CockroachDB, Redis, GraphQL, and Kubernetes.
  • Experience analyzing and working with existing codebases and collaborating across technical teams.
  • Strong communication and collaboration skills, with the ability to help other engineers build, operate, and troubleshoot reliable systems.
  • Pragmatic problem-solving mindset, with good judgment around when to improve tooling, simplify solutions, or avoid unnecessary complexity.
  • Ability to work autonomously, take ownership of projects, mentor less‑experienced engineers, and operate effectively in a fast‑moving environment.
  • A proactive, growth‑oriented approach and genuine interest in building reliable infrastructure with meaningful real‑world impact.
Benefits:
  • Remote work opportunity with flexibility and autonomy.
  • Opportunity to work on large-scale, mission‑driven financial technology infrastructure.
  • High degree of ownership across projects, from problem definition through production monitoring.
  • Professional growth through exposure to complex observability, reliability, performance, and platform engineering challenges.
  • Opportunities to mentor and collaborate with experienced engineers across distributed teams.
  • Inclusive and diverse international engineering environment.
  • Opportunity to help define the scope, standards, tooling, and practices of a growing Performance & Observability function.
  • Exposure to modern technologies including OpenTelemetry, Kubernetes, cloud infrastructure, Datadog, Honeycomb, Sentry, Pyroscope, and related observability platforms.
  • Competitive compensation and benefits package, according to local employment terms.
  • Flexible, autonomy‑focused work culture designed to support productivity and continuous learning.

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

LanceSoft, Inc. • Montreal (administrative region)

On-site
CAD 80,000 - 120,000
Software Engineer, Observability
Software Engineer, Observability

Okta • Toronto

Hybrid
CAD 110,000 - 152,000
Equity
Bonus
Health insurance
+4
Software Engineer, Observability
Software Engineer, Observability

United States Digital Space LLC • Toronto

Hybrid
CAD 110,000 - 152,000
Equity
Bonus
Health insurance
+4
Principal Observability Engineer
Principal Observability Engineer

ISG Search Inc • Toronto

On-site
CAD 150,000 - 190,000
Azure or GCP experience
Red Hat technologies
Virtualization
+3
GCP Observability Engineer
GCP Observability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 85,000 - 115,000
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

Jobgether • Ottawa, Toronto, Vancouver

Hybrid
CAD 185,000 - 245,000
Competitive salary CAD $184,500–$244,_
Equity ownership
Flexible work model
+9
Software Engineer, Data Platform
Software Engineer, Data Platform

Jobgether • Toronto

Hybrid
CAD 130,000 - 165,000
Remote work (Toronto remote)
Competitive CAD salary
Comprehensive medical benefits
+2
SRE Observability Engineer
SRE Observability Engineer

Tata Consultancy Services • Toronto

On-site
CAD 90,000 - 120,000
Senior Observability Architect
Senior Observability Architect

1050 Suncor Energy Services Inc. • Calgary

On-site
CAD 120,000 - 160,000
Strong compensation
Benefits: health, dental, vision
Generous paid time off
+1
Senior Observability Architect
Senior Observability Architect

Suncor Energy • Calgary

On-site
CAD 120,000 - 160,000
Competitive compensation
Benefits
Generous time off
+1