Site Reliability Engineer – Observability

Decskill

Portugal

Remote

EUR 55,000 - 75,000

Full time

5 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Long-term projects
Growth opportunities
People-first culture
Ownership & collaboration

Job summary

Decskill seeks a Site Reliability Engineer – Observability to design, implement, and maintain our observability stack. You will build reliable monitoring pipelines covering metrics, logs, traces, and RUM, using Grafana Cloud, Loki, Mimir, Tempo, OpenTelemetry, and related tools.

You will work with development and operations teams to promote observability by design, define best practices, and optimize costs while driving long-term platform evolution in a remote-friendly setup.

Qualifications

  • Solid experience in Site Reliability Engineering, Platform Engineering, DevOps, or Observability.
  • Hands-on experience with observability technologies such as Grafana, Loki, Mimir, Tempo, Grafana Cloud, Alloy, and/or OpenTelemetry.
  • Strong understanding of metrics, logs, distributed tracing, telemetry pipelines, and monitoring architectures.
  • Experience designing SLIs, SLOs, SLAs, dashboards, and alerting strategies.

Responsibilities

  • Collaborate with software development and operations teams to design scalable observability solutions.
  • Design and implement scalable observability pipelines and dashboards.
  • Improve reliability and usability of monitoring and telemetry.
  • Automate operational processes and reduce manual maintenance.
  • Promote observability best practices across teams.
  • Contribute to architectural discussions and platform strategy.

Skills

Site Reliability
Observability
DevOps
Cloud-native
Automation
Grafana Cloud
OpenTelemetry
SLIs/SLOs/SLAs
Monitoring architectures
Team collaboration

Tools

Grafana
Loki
Mimir
Tempo
Alloy
OpenTelemetry

Job description

We are growing. And we are looking for talented people.

At Decskill, we believe that technological excellence is driven by human talent.

We are an IT consulting company with more than 10 years of consolidated experience in the market, focused on building long-term relationships with both our clients and our people. Today, we are a community of over 800 professionals, working from Lisbon, Porto, and Madrid, contributing to impactful technology initiatives.

As part of the Astek Group, we combine a strong local culture with a global presence, being active in around 23 countries across 4 continents. This allows us to offer an international context, diverse challenges, and long-term career opportunities, while staying close to our teams.

We would like to meet a Site Reliability Engineer – Observability for a project with a remote working model.

Your Mission:
  • Design, implement, and maintain observability solutions covering metrics, logs, traces, and Real User Monitoring (RUM).
  • Work extensively with Grafana Cloud, Grafana Tempo, Loki, Mimir, Alloy, and OpenTelemetry.
  • Build reliable monitoring and alerting pipelines based on SLOs and SLAs, with a strong focus on automation and low operational overhead.
  • Ensure the health, quality, and integrity of observability data flows, from instrumentation and collection to dashboards and alerting.
  • Collaborate closely with development and operations teams to promote and implement observability by design throughout the software development lifecycle.
  • Define, document, and promote observability best practices, standards, and patterns across the organization.
  • Contribute to the modernization of our observability landscape by replacing, simplifying, and evolving legacy monitoring and alerting solutions.
  • Monitor observability-related costs and contribute to FinOps initiatives, identifying opportunities to optimize resource usage and platform spend.
  • Take ownership of technical challenges and contribute to architectural decisions and the long-term evolution of the observability platform.
The Team & Day-to-Day:

You will join a motivated, cross-functional team responsible for implementing and scaling our new observability stack.

Your day-to-day work will include:
  • Collaborating with software development and operations teams.
  • Designing and implementing scalable observability solutions.
  • Improving the reliability and usability of monitoring and telemetry.
  • Automating operational processes and reducing manual maintenance.
  • Helping teams adopt observability best practices.
  • Investigating technical challenges and driving continuous improvement.
  • Contributing to architectural discussions and platform strategy.

This is a role with strong technical ownership and influence, where your decisions will directly contribute to the reliability, scalability, and evolution of our platform.

What We’re Looking For:
  • Solid experience in Site Reliability Engineering, Platform Engineering, DevOps, or Observability.
  • Hands-on experience with observability technologies such as Grafana, Loki, Mimir, Tempo, Grafana Cloud, Alloy, and/or OpenTelemetry.
  • Strong understanding of metrics, logs, distributed tracing, telemetry pipelines, and monitoring architectures.
  • Experience designing SLIs, SLOs, SLAs, dashboards, and alerting strategies.
  • Experience with automation and infrastructure/platform engineering.
  • Strong understanding of cloud-native and distributed systems.
  • Ability to work collaboratively across development and operations teams.
  • A strong focus on reliability, scalability, automation, and operational simplicity.
  • Good communication skills and the ability to influence technical decisions across teams.

Are you looking for an environment that values curiosity and commitment? Here, your contribution has real impact and individual growth is taken seriously.

Find with us the right opportunity to grow!

What you can expect from us
  • Long-term projects with national and international context
  • Opportunities to grow technically and professionally
  • A people-first culture, focused on transparency and trust
  • Teams that value ownership, collaboration and stability

Your next challenge might start here.

At Decskill, we are committed to equal opportunities and non-discrimination. We promote a culture of diversity and inclusion, where recruitment and career progression are based solely on talent, regardless of age, gender, ethnicity, race, nationality, or any other form of discrimination incompatible with human dignity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DevOps Engineer | Cloud-Native Platform
DevOps Engineer | Cloud-Native Platform

Decskill • Lisboa

Hybrid
EUR 60,000 - 90,000
DevOps Infrastructure Engineer | SRE
DevOps Infrastructure Engineer | SRE

Decskill • Lisboa

Hybrid
EUR 55,000 - 75,000
Observability Engineer
Observability Engineer

LUZA Group • Leiria

Remote
EUR 45,000 - 65,000
Remote work when possible
Equipment provided
Benefits plan
Tech Lead de Data & Analytics
Tech Lead de Data & Analytics

Decskill • Lisboa

On-site
EUR 85,000 - 130,000
DevOps & Technical Expert
DevOps & Technical Expert

Decskill • Lisboa

On-site
EUR 60,000 - 90,000
Devops Engineer
Devops Engineer

Decskill • Lisboa

On-site
EUR 40,000 - 75,000
DevOps Infrastructure Engineer
DevOps Infrastructure Engineer

Decskill • Lisboa

On-site
EUR 55,000 - 75,000
DevOps & SRE Developer
DevOps & SRE Developer

Decskill • Lisboa

On-site
EUR 45,000 - 65,000
DEVOPS
DEVOPS

Decskill • Porto

On-site
EUR 48,000 - 72,000
DEVOPS for Application Production Support
DEVOPS for Application Production Support

Decskill • Lisboa

On-site
EUR 42,000 - 65,000