Cloud Site Reliability Engineer II

Jobvite, Inc.

Ann Arbor (MI)

On-site

USD 110,000 - 170,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity options
Health benefits
Retirement plan with employer match
Career growth opportunities
Flexible PTO
Volunteer opportunities

Job summary

Barracuda Networks is seeking a Cloud Site Reliability Engineer II to join the CloudOps Platform team. You will focus on observability, monitoring, dashboards, and automation that power Barracuda’s multi-tenant Kubernetes platform across AWS and Azure.

You'll design centralized observability stacks, create dashboards, establish reliable alerting, and automate telemetry deployment. The role emphasizes collaboration with internal teams and embracing AI-assisted workflows.

Qualifications

  • 2–4 years of experience with public cloud infrastructure (AWS and/or Azure).
  • 1–2+ years deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).
  • Hands-on experience building dashboards in Grafana and modern observability stacks.
  • Working knowledge of Terraform/Terragrunt and GitOps delivery (ArgoCD or Flux).
  • Scripting skills in Python or Bash; Go is a plus.
  • Interest in applying AI coding tools in daily engineering workflows.
  • Strong analytical troubleshooting and collaboration skills.

Responsibilities

  • Operate and scale centralized telemetry infrastructure (Loki, Mimir, Tempo, Grafana).
  • Design Grafana dashboards and health overviews for services and tenants.
  • Establish reliable alerting, SLO/SLI tracking, and alert routing.
  • Automate deployment of log collectors, metric exporters, and monitoring agents across multi-cluster environments using GitOps.
  • Contribute to Kubernetes platform health and infrastructure modernization.
  • Leverage AI coding tools to build automation and scripts.
  • Collaborate with product and tenant teams on observability onboarding.

Skills

Observability
Kubernetes
Grafana
Prometheus
Terragrunt
Terraform
ArgoCD
GitHub Actions
AWS
Azure
Python
Bash
Go
Claude Code
OpenCode
Codex CLI
GitHub Copilot

Tools

Kubernetes
Grafana
Prometheus
Terraform
Terragrunt
ArgoCD
GitHub Actions
AWS
Azure
Python
Bash
Go

Job description

Come join our passionate team! Barracuda is a leading cybersecurity company providing complete protection against complex threats. Our platform protects email, data, applications, and networks with innovative solutions, and a managed XDR service, to strengthen cyber resilience. Hundreds of thousands of IT professionals and managed service providers worldwide trust us to protect and support them with solutions that are easy to buy, deploy, and use.

We know a diverse workforce adds to our collective value and strength as an organization. Barracuda Networks is proud to be an Equal Opportunity Employer, committed to equal employment opportunity and equitable compensation regardless of race, gender, religion, sex, sexual orientation, national origin, or disability.

Envision yourself at Barracuda

As a Cloud Site Reliability Engineer II on the CloudOps Platform team, you will focus on the observability, monitoring, dashboards, and automation systems that power Barracuda’s next-generation multi-Tenant Kubernetes platform. While your primary mission centers on delivering deep operational visibility, reliable telemetry pipelines, and proactive alerting, you will also play a key role in improving the broader Kubernetes platform running across AWS and Azure.

Our team values collaborative knowledge sharing, automated reliability, and modern engineering practices. We actively embrace AI-assisted workflows (Claude Code, OpenCode, Codex CLI) to accelerate development, diagnostics, and routine platform maintenance. In this role, you will build and operate centralized observability stacks (Grafana, Loki, Mimir, Tempo), create actionable dashboards, automate telemetry via GitOps and Terragrunt, and partner with internal engineering teams to optimize application reliability in production.

Tech Stack Exposure:
  • Observability & Monitoring: Grafana, Prometheus / Mimir, Loki, Tempo, OpenTelemetry / Grafana Alloy, Sloth (SLOs), Alertmanager
  • Orchestration & Compute: Kubernetes (AWS EKS, Azure AKS), Helm, Kustomize
  • IaC & Cloud Provisioning: Terragrunt, Terraform
  • GitOps & CI/CD: ArgoCD, GitHub Actions
  • Public Clouds: Amazon Web Services (AWS), Microsoft Azure
  • Languages & Scripting: Python, Bash (Go is a plus)
  • AI Developer Tooling: Claude Code, OpenCode, Codex CLI, GitHub Copilot
What you’ll be working on
  • Operating, scaling, and automating our centralized LGTM telemetry infrastructure (Loki for logs, Mimir/Prometheus for metrics, Tempo for distributed tracing, and Grafana for unified visualization).
  • Designing intuitive, high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
  • Establishing reliable alerting strategies, SLO/SLI tracking (via Sloth), and notification routing to detect and resolve degradation before it impacts production systems.
  • Automating the deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS and AKS environments using GitOps (ArgoCD) and Terragrunt.
  • Directly contributing to core MTK Kubernetes platform health, performance tuning, and infrastructure modernization.
  • Leveraging modern AI coding tools (Claude Code, OpenCode, Codex CLI) to build automation, diagnostic tooling, and operational scripts.
  • Collaborating closely with internal product and tenant teams to assist with observability onboarding, distributed tracing instrumentation, and performance troubleshooting.
What you bring to the role
  • 2–4 years of experience working with public cloud infrastructure (AWS and/or Azure) with a strong passion for observability, monitoring, and systems reliability.
  • 1–2+ years of hands‑on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).
  • Practical experience configuring, operating, or building dashboards in modern observability stacks (Grafana, ELK, Splunk, etc).
  • Working knowledge of Infrastructure as Code using Terraform and/or Terragrunt, and GitOps delivery workflows (ArgoCD or Flux).
  • Solid scripting skills in Python or Bash for system automation, telemetry pipelines, and operational tooling (Go is a plus).
  • Curiosity and eagerness to leverage AI coding agents (Claude Code, OpenCode, Codex CLI, GitHub Copilot) in everyday engineering workflows.
  • Strong analytical troubleshooting instincts, clear communication skills, and a collaborative team mindset.
What you’ll get from us

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda.

  • Equity, in the form of non-qualifying options
  • High-quality health benefits
  • Retirement Plan with employer match
  • Career-growth opportunities
  • Flexible Time Off and Paid Time Off benefits
  • Volunteer opportunities

Job ID - 27-0459

#LI-hybrid

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

Barracuda Networks Inc. • Ann Arbor (MI)

Hybrid
USD 120,000 - 160,000
Equity options
Health benefits
Retirement plan with employer match
+3
Senior Software Engineer
Senior Software Engineer

Barracuda Networks Inc. • Alpharetta (GA)

Hybrid
USD 120,000 - 160,000
Equity options
High-quality health benefits
Retirement plan with employer match
+3
Senior Software Engineer
Senior Software Engineer

Barracuda • Ann Arbor (MI)

On-site
USD 120,000 - 180,000
Equity options
Health benefits
401(k) matching employer contribution
+3
Senior Software Engineer
Senior Software Engineer

Barracuda • Alpharetta (GA)

On-site
USD 130,000 - 180,000
Equity in non-qualifying options
High-quality health benefits
Retirement Plan with employer match
+3
Manager, Offensive Security
Manager, Offensive Security

Barracuda • South Carolina

On-site
USD 120,000 - 180,000
Equity (non-qualifying options)
High-quality health benefits
Retirement Plan with employer match
+2
Manager, Offensive Security
Manager, Offensive Security

Barracuda • Town of Florida (NY)

On-site
USD 150,000 - 230,000
Equity options
Health benefits
Retirement plan
+3
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Principal Site Reliability Engineer / DevOps (Prisma AIRS)
Principal Site Reliability Engineer / DevOps (Prisma AIRS)

Jobs Paloaltonetworks • California (MO)

On-site
USD 184,000 - 297,000
Senior SRE - Government Cloud Operations
Senior SRE - Government Cloud Operations

Cato Networks • United States

On-site
USD 140,000 - 200,000
Health insurance
401(k)
Stock options
+4
Platform Architect
Platform Architect

Biorce • Austin (TX)

On-site
USD 130,000 - 160,000
Company-sponsored premium gym membership
Modern equipment and productivity tooling
Regular company events and team offsites