Site Reliability Engineer, Observability & Alerting

Cobira

Denmark

Remote

DKK 700,000 - 800,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cobira is seeking a mid-level Site Reliability Engineer to own observability and alerting across our platform. You will design alerting architectures, implement dashboards, and reduce incident noise while collaborating with platform engineers shipping the product.

You'll build tooling in Python to automate tasks, participate in on-call rotations, and contribute to infrastructure work. The role is based in Copenhagen with flexible hours and an autonomous, hands-on team culture.

Qualifications

  • 3–5 years of experience in SRE, platform engineering, or a strong DevOps role.
  • Hands-on experience building observability systems (alerting, dashboards, tracing, logging) - New Relic, Datadog, Grafana, or similar.
  • Solid Linux fundamentals and comfort operating in cloud-hosted VM environments.
  • Experience with containerised workloads (Docker, Docker Compose).
  • A systematic approach to debugging - you form hypotheses, isolate variables, and document what you find.
  • Good written communication; we write things down.
  • Nice-to-have: Kafka or other message streaming systems; Redis or caching technologies; network infrastructure familiarity; IoT/connectivity exposure; NRQL or query language exposure; IaC (Terraform/Ansible); Kubernetes familiarity.

Responsibilities

  • Own and mature our observability stack — alerting policies, dashboards, on-call runbooks, and incident response workflows.
  • Design and tune alert conditions in New Relic (NRQL, baseline/anomaly detection, composite conditions) to minimize noise and maximize signal.
  • Identify gaps in our monitoring coverage across services, message queues, infrastructure, and network links.
  • Build and maintain tooling that helps the team understand system behavior — not just when things break, but before they do.
  • Scripting ability in Python - enough to automate, glue systems together, and write a useful tool when one doesn't exist.
  • Collaborate with platform engineers on SLIs, SLOs, and error budgets.
  • Participate in on-call rotation and drive post-incident improvements.
  • Contribute to infrastructure work when needed

Skills

SRE
Observability
Docker
Python scripting
Cloud VM

Tools

New Relic
Datadog
Grafana
NRQL

Job description

Site Reliability Engineer, Observability & Alerting
The Role

We're looking for a mid-level SRE to own observability and alerting across our platform. Right now our monitoring lives primarily in New Relic, and while we've built solid foundations - Kafka consumer lag tracking, infrastructure health dashboards, custom NRQL alert policies - we know there's a lot more to do. You'll be the person driving this forward.

This isn't a pure ops role. You'll write code, design alerting architectures, and work closely with the engineers shipping the platform. When something is on fire, you'll be one of the people who actually understands why.

What You'll Do

Own and mature our observability stack — alerting policies, dashboards, on-call runbooks, and incident response workflows

Design and tune alert conditions in New Relic (NRQL, baseline/anomaly detection, composite conditions) to minimize noise and maximize signal

Identify gaps in our monitoring coverage across services, message queues, infrastructure, and network links

Build and maintain tooling that helps the team understand system behavior — not just when things break, but before they do

Scripting ability in Python - enough to automate, glue systems together, and write a useful tool when one doesn't exist

Collaborate with platform engineers on SLIs, SLOs, and error budgets

Participate in on-call rotation and drive post-incident improvements

Contribute to infrastructure work when needed

What We’re Looking For
Must-have:

3–5 years of experience in SRE, platform engineering, or a strong DevOps role

Hands-on experience building and maintaining observability systems (alerting, dashboards, tracing, logging) - New Relic, Datadog, Grafana, or similar

Solid Linux fundamentals and comfort operating in cloud-hosted VM environments

Experience with containerised workloads (Docker, Docker Compose)

A systematic approach to debugging - you form hypotheses, isolate variables, and document what you find

Good written communication; we write things down

Nice-to-have:

Experience with Kafka or other message streaming systems

Experience with Redis or other caching technologies

Familiarity with network-level infrastructure (VPNs, firewall rules, routing)

Exposure to telecom or IoT connectivity domains

Experience with NRQL or another query language for observability platforms

IaC experience (Terraform, Ansible, or similar)

Familiarity with Kubernetes - we're not there yet, but directionally heading that way

What We Offer

A technically honest environment - we'll tell you what's messy and where improvement is needed

Meaningful ownership from day one; no layers of process between you and the problem

A compact, experienced team in Copenhagen

Competitive salary based on experience

Flexible hours and autonomy over how you work, within an on-site team culture

The chance to shape the reliability culture of a growing IoT and connectivity platform

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Trackunit • Aarhus

On-site
DKK 700,000 - 900,000
Flexible and hybrid working
Training & development
Site Reliability Engineer
Site Reliability Engineer

Trackunit • Kolding

On-site
DKK 700,000 - 900,000
Flexible setup
Remote-friendly
Professional development
+1
Site Reliability Engineer
Site Reliability Engineer

Trackunit • Aalborg

Hybrid
DKK 560,000 - 920,000
Senior DevOps / Observability Engineer – AI Runtime & Platform Monitoring
Senior DevOps / Observability Engineer – AI Runtime & Platform Monitoring

Twoday • Københavns Kommune

On-site
DKK 700,000 - 900,000
Pension
Health insurance
Professional development opportunities
Site Reliability Engineer Engineering · ·
Site Reliability Engineer Engineering · ·

Trackunit • København, Kolding, Aarhus, Aalborg

On-site
DKK 650,000 - 850,000
SRE: Observability & Alerting for IoT - Flexible Hours
SRE: Observability & Alerting for IoT - Flexible Hours

Cobira • Denmark

On-site
DKK 700,000 - 800,000
SRE & Platform Engineer
SRE & Platform Engineer

Nnit A/S • København

Hybrid
DKK 900,000 - 1,300,000
Competitive salary
Remote work options
Career development
+1
Senior Observability Engineer
Senior Observability Engineer

Saxo Bank A/S • Gentofte Kommune

On-site
DKK 900,000 - 1,200,000
Site Reliability Engineer - Trackunit
Site Reliability Engineer - Trackunit

Trackunit A/S • Aalborg

On-site
DKK 700,000 - 900,000
Senior Network Engineer (Linux)
Senior Network Engineer (Linux)

Keepit • København

Hybrid
DKK 650,000 - 950,000
Competitive salary
Pension scheme
Hybrid work model
+3