Cloud Site Reliability Engineer II

Barracuda

Ottawa

On-site

CAD 100,000 - 120,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity options
Health benefits
Retirement plan with employer match
Career growth opportunities
Flexible time off

Job summary

Barracuda is seeking a Cloud Site Reliability Engineer II to join the CloudOps Platform team in Ottawa. You will focus on observability, monitoring, dashboards, and automation for a multi-tenant Kubernetes platform across AWS and Azure.

You’ll build centralized telemetry stacks, design dashboards, and implement automated telemetry pipelines with GitOps, Terragrunt, and modern AI tooling. Collaboration with product teams is essential for observability onboarding and performance tuning.

Qualifications

  • 2-4 years of public cloud infrastructure experience (AWS/Azure).
  • 1-2+ years deploying or operating containerized workloads in Kubernetes (EKS/AKS).
  • Experience configuring dashboards in Grafana/ELK/Splunk stacks.
  • IaC with Terraform/Terragrunt and GitOps workflows (ArgoCD/Flux).
  • Scripting skills in Python or Bash for automation.

Responsibilities

  • Operate, scale, and automate centralized LGTM telemetry infrastructure (Loki, Mimir/Prometheus, Tempo, Grafana).
  • Design intuitive Grafana dashboards and health overviews for platform services and tenant workloads.
  • Establish reliable alerting and SLO/SLI tracking with proactive degradation detection.
  • Automate deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS/AKS with GitOps (ArgoCD) and Terragrunt.
  • Contribute to MTK Kubernetes platform health and modernization.
  • Leverage AI coding tools to build automation and diagnostics.
  • Collaborate with product/tenant teams on observability onboarding and performance troubleshooting.

Skills

Kubernetes
Observability
Grafana
Prometheus
Terragrunt
Terraform
ArgoCD
Python
GitOps
Cloud infrastructure

Tools

Grafana
Prometheus
Mimir
Tempo
Loki
OpenTelemetry
Terragrunt
Terraform
ArgoCD
Kubernetes

Job description

Come join our passionate team! Barracuda is a leading cybersecurity company providing complete protection against complex threats. Our platform protects email, data, applications, and networks with innovative solutions, and a managed XDR service, to strengthen cyber resilience. Hundreds of thousands of IT professionals and managed service providers worldwide trust us to protect and support them with solutions that are easy to buy, deploy, and use.
We are committed to a candidate selection process and work environment that is inclusive and barrier free. To ensure candidates are assessed in a fair and equitable manner, accommodations will be provided to prospective employees in accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and the Ontario Human Rights Code.

Envision yourself at Barracuda

As a Cloud Site Reliability Engineer II on the CloudOps Platform team, you will focus on the observability, monitoring, dashboards, and automation systems that power Barracuda’s next-generation multi-Tenant Kubernetes platform. While your primary mission centers on delivering deep operational visibility, reliable telemetry pipelines, and proactive alerting, you will also play a key role in improving the broader Kubernetes platform running across AWS and Azure.

Our team values collaborative knowledge sharing, automated reliability, and modern engineering practices. We actively embrace AI-assisted workflows (Claude Code, OpenCode, Codex CLI) to accelerate development, diagnostics, and routine platform maintenance. In this role, you will build and operate centralized observability stacks (Grafana, Loki, Mimir, Tempo), create actionable dashboards, automate telemetry via GitOps and Terragrunt, and partner with internal engineering teams to optimize application reliability in production.

Tech Stack Exposure
  • Observability & Monitoring: Grafana, Prometheus / Mimir, Loki, Tempo, OpenTelemetry / Grafana Alloy, Sloth (SLOs), Alertmanager
  • Orchestration & Compute: Kubernetes (AWS EKS, Azure AKS), Helm, Kustomize
  • IaC & Cloud Provisioning: Terragrunt, Terraform
  • GitOps & CI/CD: ArgoCD, GitHub Actions
  • Public Clouds: Amazon Web Services (AWS), Microsoft Azure
  • Languages & Scripting: Python, Bash (Go is a plus)
  • AI Developer Tooling: Claude Code, OpenCode, Codex CLI, GitHub Copilot
What You’ll Be Working On
  • Operating, scaling, and automating our centralized LGTM telemetry infrastructure (Loki for logs, Mimir/Prometheus for metrics, Tempo for distributed tracing, and Grafana for unified visualization).
  • Designing intuitive, high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
  • Establishing reliable alerting strategies, SLO/SLI tracking (via Sloth), and notification routing to detect and resolve degradation before it impacts production systems.
  • Automating the deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS and AKS environments using GitOps (ArgoCD) and Terragrunt.
  • Directly contributing to core MTK Kubernetes platform health, performance tuning, and infrastructure modernization.
  • Leveraging modern AI coding tools (Claude Code, OpenCode, Codex CLI) to build automation, diagnostic tooling, and operational scripts.
  • Collaborating closely with internal product and tenant teams to assist with observability onboarding, distributed tracing instrumentation, and performance troubleshooting.
What You Bring To The Role
  • 2-4 years of experience working with public cloud infrastructure (AWS and/or Azure) with a strong passion for observability, monitoring, and systems reliability.
  • 1-2+ years of hands-on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).
  • Practical experience configuring, operating, or building dashboards in modern observability stacks (Grafana, ELK, Splunk, etc).
  • Working knowledge of Infrastructure as Code using Terraform and/or Terragrunt, and GitOps delivery workflows (ArgoCD or Flux).
  • Solid scripting skills in Python or Bash for system automation, telemetry pipelines, and operational tooling (Go is a plus).
  • Curiosity and eagerness to leverage AI coding agents (Claude Code, OpenCode, Codex CLI, GitHub Copilot) in everyday engineering workflows.
  • Strong analytical troubleshooting instincts, clear communication skills, and a collaborative team mindset.
What You’ll Get From Us

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda.

  • Equity, in the form of non-qualifying options
  • High-quality health benefits
  • Retirement Plan with employer match
  • Career-growth opportunities
  • Flexible Time Off and Paid Time Off benefits
  • Volunteer opportunities
What You’ll Get From Us

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda. In addition, you will receive equity, in the form of non-qualifying options.

The anticipated salary range for this role is $100,000 CAD to $120,000 CAD. Actual compensation offered will be dependent upon the individual's skills, experience, and qualifications as they directly relate to the requirements of the position, the budget for the position, and applicable employment law.

Location: Ottawa, ON

27-0459(a)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Senior Engineer
Cloud Site Reliability Senior Engineer

Jobvite, Inc. • Ottawa

On-site
CAD 110,000 - 140,000
Equity options (non-qualifying)
Cloud Site Reliability Senior Engineer
Cloud Site Reliability Senior Engineer

barracuda-networks-inc • Ottawa

Hybrid
CAD 110,000 - 140,000
Equity options
Internal mobility
Cross training opportunities
Manager, Cloud Services and Site Reliability
Manager, Cloud Services and Site Reliability

Barracuda • Ottawa

On-site
CAD 136,000 - 182,000
Equity options
Health benefits
Retirement plan
+3
Software Development Engineer in Test II
Software Development Engineer in Test II

Barracuda • Ottawa

On-site
CAD 74,000 - 98,000
Equity options
Internal mobility
Career progression
Technical Support Representative
Technical Support Representative

Barracuda • Ottawa

Hybrid
CAD 43,000 - 58,000
Equity options
Health benefits
Retirement plan
+3
Cloud SRE II: Observability, K8s & Equity Options
Cloud SRE II: Observability, K8s & Equity Options

Barracuda • Ottawa

On-site
CAD 100,000 - 120,000
Equity options
Health benefits
Retirement plan with employer match
+2
Senior Software Developer
Senior Software Developer

BlueCat • Toronto

On-site
CAD 130,000 - 150,000
Professional Development Budget
Dedicated Wellness Days and Wellness W
Lifestyle Spending Account
+1
Partner Development Manager II
Partner Development Manager II

Barracuda • Ottawa

On-site
CAD 106,000 - 142,000
Equity options
Health benefits
Retirement plan with employer match
+3
Senior Platform Systems Engineer
Senior Platform Systems Engineer

Bettermode • Toronto

On-site
CAD 160,000 - 180,000
Health benefits
Unlimited vacation
Parental leave
+5
Cloud Engineer 2, Site Reliability Engineering
Cloud Engineer 2, Site Reliability Engineering

Kinaxis • Ottawa, Toronto

Hybrid
CAD 95,000 - 130,000
Flexible vacation & Kinaxis Days
Flexible work options
Well-being programs
+1