Cloud SRE II - Observability & Automation for Kubernetes

barracuda-networks-inc

Ottawa

Hybrid

CAD 100,000 - 120,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Barracuda is looking for a Cloud Site Reliability Engineer II in Ottawa to strengthen our multi-tenant Kubernetes platform. You will own observability, telemetry pipelines, dashboards, and automation across AWS and Azure, delivering proactive reliability improvements.

You will build centralized stacks (Grafana, Prometheus, Loki, Tempo) and implement GitOps with ArgoCD and Terragrunt, partnering with product teams to enhance production performance and resiliency.

Qualifications

  • 2–4 years of public cloud and observability experience.
  • 1–2+ years deploying or operating containerized workloads in Kubernetes (EKS/AKS).
  • Strong scripting in Python or Bash; Go is a plus.

Responsibilities

  • Operate, scale, and automate centralized telemetry infrastructure (Loki, Mimir, Tempo) and Grafana dashboards.
  • Design intuitive dashboards and health overviews for platform services and tenant workloads.
  • Establish reliable alerting and SLO/SLI tracking with clear notifications.
  • Automate deployment of log collectors and metrics exporters across multi-cluster EKS/AKS using GitOps (ArgoCD) and Terragrunt.
  • Contribute to Kubernetes platform health, performance tuning, and infra modernization.
  • Leverage AI coding tools to build automation and diagnostic tooling.
  • Collaborate with product and tenant teams on observability onboarding and tracing instrumentation.

Skills

Observability
Kubernetes
Grafana
Terraform
Terragrunt
Python
Bash
Go
AWS
Azure
GitOps
AI tooling

Tools

Grafana
Prometheus
Mimir
Loki
Tempo
OpenTelemetry
ArgoCD
GitHub Actions
Terraform
Terragrunt
AWS
Azure

Job description

Barracuda is looking for a Cloud Site Reliability Engineer II in Ottawa to strengthen our multi-tenant Kubernetes platform. You will own observability, telemetry pipelines, dashboards, and automation across AWS and Azure, delivering proactive reliability improvements.

You will build centralized stacks (Grafana, Prometheus, Loki, Tempo) and implement GitOps with ArgoCD and Terragrunt, partnering with product teams to enhance production performance and resiliency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer II
Cloud Site Reliability Engineer II

barracuda-networks-inc • Ottawa

Hybrid
CAD 100,000 - 120,000
Cloud Site Reliability Engineer II
Cloud Site Reliability Engineer II

Jobvite, Inc. • Ottawa

Hybrid
CAD 100,000 - 120,000
Equity options
Health benefits
Retirement plan with employer match
+3
SRE Manager — Lead Cloud Reliability & Automation
SRE Manager — Lead Cloud Reliability & Automation

Barracuda Networks, Inc. • Ottawa

Hybrid
CAD 136,000 - 182,000
Non-qualifying options equity
High-quality health benefits
Retirement plan with employer match
+3
Senior SRE: Kubernetes Reliability for Cloud UI Services
Senior SRE: Kubernetes Reliability for Cloud UI Services

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Senior SRE: Cloud Reliability & Observability
Senior SRE: Cloud Reliability & Observability

Redwood Software • Markham

On-site
CAD 125,000 - 145,000
Senior Cloud Reliability Engineer (AWS/Kubernetes)
Senior Cloud Reliability Engineer (AWS/Kubernetes)

Enverus • Calgary

On-site
CAD 120,000 - 180,000
Founding SRE: Cloud Reliability & Platform Lead
Founding SRE: Cloud Reliability & Platform Lead

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Manager, Cloud Services and Site Reliability
Manager, Cloud Services and Site Reliability

Barracuda Networks, Inc. • Ottawa

Hybrid
CAD 136,000 - 182,000
Non-qualifying options equity
High-quality health benefits
Retirement plan with employer match
+3
Manager, Cloud Services and Site Reliability
Manager, Cloud Services and Site Reliability

Barracuda • Ottawa

On-site
CAD 136,000 - 182,000
Equity options
Health benefits
Retirement plan
+3
Cloud Site Reliability Engineer (SRE) — AWS, Kubernetes
Cloud Site Reliability Engineer (SRE) — AWS, Kubernetes

OpenText • Waterloo

On-site
CAD 70,000 - 100,000