Senior Site Reliability Engineer

Claranet Portugal

Portugal

Presencial

EUR 45 000 - 65 000

Tempo integral

há 38 horas
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.

Ultrapassa os filtros ATS

Resumo da oferta

Claranet Portugal is seeking a Site Reliability Engineer to join our team and help us build and operate reliable, scalable, and secure cloud platforms. The role blends operational excellence with engineering projects across Azure and Kubernetes environments.

You will design automation, manage Terraform IaC, improve CI/CD, and implement observability with Datadog, Prometheus, and Grafana, collaborating with development and infrastructure teams to deliver resilient services.

Qualificações

  • Degree in Computer Science, Engineering or related field.
  • At least 4 years of experience in SRE, DevOps, Platform Engineering or similar.
  • Hands-on experience with Microsoft Azure, particularly compute, networking and storage services.
  • Practical experience with Kubernetes; AKS is an advantage.
  • Experience with Terraform or another Infrastructure as Code tool.
  • Familiarity with CI/CD practices and version control systems.
  • Experience with monitoring, logging and alerting platforms such as Datadog, Azure Monitor, Prometheus, Grafana or equivalent.
  • Good scripting skills in Bash, Python or PowerShell.
  • Understanding of software development and deployment practices.
  • Experience with .NET and/or Java, microservices or business applications deployed on Kubernetes is a strong advantage.
  • Ability to troubleshoot complex technical issues in a structured and collaborative way.
  • Good written and verbal communication skills in English.

Responsabilidades

  • Monitoring and maintaining cloud and Kubernetes platforms to ensure high availability and performance.
  • Investigating and resolving incidents, conducting root cause analysis, and driving continuous service improvements.
  • Designing, deploying, and managing scalable infrastructure in Azure.
  • Managing Kubernetes environments, preferably with AKS (Azure Kubernetes Service).
  • Developing and maintaining Infrastructure as Code using Terraform.
  • Building and improving CI/CD pipelines and automating operational processes.
  • Implementing observability solutions, including monitoring, logging, tracing, and alerting tools such as Datadog.
  • Defining and tracking reliability and performance metrics (SLIs, SLOs, and error budgets).
  • Collaborating with development and infrastructure teams to deliver reliable, secure, and maintainable platform solutions.
  • Promoting DevOps, automation, knowledge sharing, and a culture of continuous improvement.

Conhecimentos

SRE
DevOps
CI/CD
Observability
Automation
Infrastructure as Code
Kubernetes
Azure
Scripting
Datadog
English

Formação académica

Degree in Computer Science, Engineering or related field

Ferramentas

AKS
Terraform
Datadog
Azure Monitor
Prometheus
Grafana
Git

Descrição da oferta de emprego

We're fast learners, hard workers, natural collaborators... and we Make Modern Happen!

Our ambition is to unlock the potential of our digital world so that organisations everywhere can innovate and thrive securely.

We aim to achieve this goal by bringing together the world’s most talented people and the most powerful technologies, combining them to address our customers' challenges and to build something stronger together.

We are looking for a Site Reliability Engineer to join our team and help us build and operate reliable, scalable and secure technology platforms.

This role combines two complementary areas of work:

  • 50% Operational Excellence: ensuring the smooth operation of our platforms, responding to service requests and incidents, troubleshooting issues and continuously improving reliability.
  • 50% Engineering & Improvement Projects: designing and implementing automation, observability, infrastructure and platform improvements that make our services more resilient and easier to operate.

This role is responsible for ensuring the reliability, performance, security, and scalability of cloud-based platforms, primarily in Azure and Kubernetes environments. The position combines operational support, infrastructure engineering, automation, and Site Reliability Engineering (SRE) practices.

Your responsibilities include:
  • Monitoring and maintaining cloud and Kubernetes platforms to ensure high availability and performance.
  • Investigating and resolving incidents, conducting root cause analysis, and driving continuous service improvements.
  • Designing, deploying, and managing scalable infrastructure in Azure.
  • Managing Kubernetes environments, preferably with AKS (Azure Kubernetes Service).
  • Developing and maintaining Infrastructure as Code using Terraform.
  • Building and improving CI/CD pipelines and automating operational processes.
  • Implementing observability solutions, including monitoring, logging, tracing, and alerting tools such as Datadog.
  • Defining and tracking reliability and performance metrics (SLIs, SLOs, and error budgets).
  • Collaborating with development and infrastructure teams to deliver reliable, secure, and maintainable platform solutions.
  • Promoting DevOps, automation, knowledge sharing, and a culture of continuous improvement.
You must have:
  • Degree in Computer Science, Engineering or a related field, or equivalent practical experience.
  • At least 4 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering or a similar role.
  • Hands-on experience with Microsoft Azure, particularly compute, networking and storage services.
  • Practical experience with Kubernetes; experience with AKS is an advantage.
  • Experience with Terraform or another Infrastructure as Code tool.
  • Familiarity with CI/CD practices and version control systems.
  • Experience with monitoring, logging and alerting platforms such as Datadog, Azure Monitor, Prometheus, Grafana or equivalent.
  • Good scripting skills in Bash, Python or PowerShell.
  • Understanding of software development and deployment practices.
  • Experience with .NET and/or Java, microservices or business applications deployed on Kubernetes is a strong advantage.
  • Ability to troubleshoot complex technical issues in a structured and collaborative way.
  • Good written and verbal communication skills in English.
We value:
  • Experience with AWS or Google Cloud.
  • Experience with .NET or Java application development.
  • Knowledge of SLI/SLO frameworks, error budgets and incident management practices.
  • Experience with distributed systems, APIs and cloud-native architectures.
  • Familiarity with security, networking and identity concepts in Azure and Kubernetes.
  • Relevant certifications, such as:
  • Microsoft Certified: Azure Solutions Architect Expert
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Site Reliability Engineer
Site Reliability Engineer

La Redoute • Leiria

Presencial
EUR 55 000 - 75 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Arcesium • Lisboa

Presencial
EUR 60 000 - 90 000
Site Reliability Engineer /Hybrid
Site Reliability Engineer /Hybrid

A2IT Technology • Leiria

Presencial
EUR 65 000 - 90 000
Site Reliability Engineer
Site Reliability Engineer

Komodo Consulting • Lisboa

Presencial
EUR 65 000 - 90 000
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

La Redoute • Viseu

Presencial
EUR 60 000 - 86 000
05 - Site Reliability Engineer
05 - Site Reliability Engineer

PPM Coachers • Lisboa

Híbrido
EUR 60 000 - 90 000
DevOps / Site Reliability Engineer | Cloud | Kubernetes | Automation
DevOps / Site Reliability Engineer | Cloud | Kubernetes | Automation

BindWorks • Porto

Presencial
EUR 45 000 - 75 000
Flexible working options
Continuous learning & certification
Long-term career development
+1
Site Reliability Engineering Manager (Data Infra)
Site Reliability Engineering Manager (Data Infra)

Complyadvantage • Lisboa

Híbrido
EUR 86 000 - 96 000
Equity participation
Private medical insurance
Unlimited Time Off Policy
+2
Site Reliability Engineer
Site Reliability Engineer

TEKEVER • Lisboa

Presencial
EUR 60 000 - 90 000
Excellent work environment
Flexible work arrangements
Professional development opportunities
+1
DevOps & Site Reliability Engineer (SRE)
DevOps & Site Reliability Engineer (SRE)

Biopharma Careers • Algés

Presencial
EUR 55 000 - 90 000