Senior Site Reliability Engineer (Remote)

Pragmatike

Lisboa

Presencial

EUR 70 000 - 110 000

Tempo integral

Há 6 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Recebe uma resposta deste empregador — um currículo e uma carta de apresentação adaptados exatamente ao que estão a contratar.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Remote work
Flexible hours
Autonomy
International team
Reliability and automation focus

Resumo da oferta

Pragmatike is hiring for a Kubernetes-focused DevOps/SRE role. You will operate Linux-based infrastructure, deploy and scale Kubernetes clusters, and own automation, observability, and incident response across multi-site environments.

The ideal candidate has expert Kubernetes in production, strong networking and Linux skills, and experience with OpenStack/Proxmox/VMware. This is a fully remote EU-timezone position with flexible hours.

Qualificações

  • Expert-level, hands-on experience operating Kubernetes in production environments.
  • Strong network engineering skills (VLANs, L2/L3 routing, VPNs, multi-site connectivity).
  • Solid Linux systems administration (Debian/Ubuntu).
  • Experience with automation workflows (Ansible, Bash/Python, Git-based).
  • Experience with observability stacks (Prometheus, Grafana, ELK, Loki, Graylog).
  • Background with virtualization tech (OpenStack, Proxmox, VMware).
  • Experience with bare-metal provisioning and MAAS.
  • Strong understanding of distributed systems and container orchestration.
  • Ability to develop SOPs and operational procedures from scratch.
  • Experience with incident response and on-call rotations.

Responsabilidades

  • Operate and maintain Linux-based infrastructure (Debian/Ubuntu).
  • Deploy and scale Kubernetes clusters across multi-site environments.
  • Oversee full cluster lifecycle: upgrades, node pools, networking, storage, security hardening.
  • Automate provisioning and operations with Ansible, Bash/Python, GitOps.
  • Design networking architecture including VLANs, L2/L3 routing, VPNs, multi-site connectivity.
  • Build automated deployment workflows (PXE, cloud-init).
  • Maintain observability stacks (Prometheus/Grafana, Loki, ELK).
  • Lead incident response and escalation across the platform.
  • Improve system availability and reduce latency at multiple levels.
  • Define and implement SLOs/SLIs; optimize alerting/monitoring pipelines.
  • Establish on-call schedules; coordinate across timezones.
  • Develop SOPs for repeatable operations and maintenance tasks.
  • Coordinate hardware maintenance for Policlouds; manage virtualization layers.
  • Plan resources for future initiatives; collaborate with dev teams.

Conhecimentos

Kubernetes production
Linux administration
Networking
Automation
Observability
Incident response
SRE practices
Autonomous work

Ferramentas

Kubernetes
OpenStack
Proxmox
VMware
MAAS
Ansible
Python/Bash
GitOps
Prometheus/Grafana

Descrição da oferta de emprego

Job Description
Job Description

Location: Fully remote EU timezone (CET ±2h)

Start date: ASAP

Languages: Fluent English is mandatory

Industry: Cloud Computing

We are hiring at Pragmatike to expand our team and drive the growth of our internal projects.

Our focus is on developing cutting-edge solutions in Cloud Computing, while fostering a culture of collaboration and innovation. Joining us means being part of a passionate team where your ideas and skills directly contribute to shaping tomorrows technologies.

If you're excited about working on ambitious projects in a dynamic and flexible environment, we'd love to hear from you!

Responsibilities
  • Operate and maintain Linux-based infrastructure (Debian/Ubuntu).
  • Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.
  • Oversee full cluster lifecycle: upgrades, node pools, networking, storage, and security hardening.
  • Implement automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows.
  • Design and maintain networking architecture including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Build automated deployment workflows (PXE boot, Preseed, cloud-init).
  • Deploy and maintain observability stacks (Prometheus/Grafana, Loki, ELK, Graylog).
  • Lead incident response and escalation activities across the platform.
  • Improve system availability and reduce latency at all levels.
  • Define and implement SLOs/SLIs at multiple infrastructure levels (physical network/hardware, platform virtualization, software services).
  • Optimize alerting and monitoring pipelines to provide actionable insights.
  • Establish and maintain on-call schedules to ensure coverage across timezones.
  • Develop Standard Operating Procedures (SOPs) for repeatable operations and maintenance tasks.
  • Coordinate physical maintenance for Policlouds (periodic maintenance, hardware issues, DC-Ops).
  • Manage virtualization and orchestration layers (OpenStack, Proxmox, VMware).
  • Help develop and maintain overall architecture across all products.
  • Plan resources for future initiatives, accounting for demand and growth projections.
  • Work with development teams to improve overall quality and optimize resource utilization.
  • Collaborate with cross-functional stakeholders (Hivenet, Policloud, Customer Success teams).
Requirements
  • Expert-level, hands-on experience operating Kubernetes in production environments.
  • Strong network engineering skills (VLANs, L2/L3 routing, VPNs, multi-site connectivity) - this is essential for the role.
  • Strong proficiency with Linux systems administration (Debian/Ubuntu).
  • Solid understanding of networking fundamentals and ability to design complex network architectures.
  • Experience building and maintaining automation workflows (Ansible, Bash/Python, Git-based).
  • Experience with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Background with virtualization technologies (OpenStack, Proxmox, VMware).
  • Experience with bare-metal provisioning and MAAS (Metal as a Service).
  • Strong understanding of distributed systems and container orchestration.
  • Process-oriented mindset with ability to develop SOPs and operational procedures from scratch.
  • Experience with incident response, escalation procedures, and on-call rotations.
  • Ability to work autonomously in a fast-paced, engineering-driven environment.
  • Strong technical skills combined with alignment to team values.
Nice To Have
  • Experience with service mesh (Istio, Linkerd) or advanced CNI implementations.
  • Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations.
  • Experience with GPU infrastructure, node preparation, or resource scheduling.
  • Familiarity with security best practices (RBAC, firewalls, network policies).
  • Exposure to IT asset management or license tracking workflows.
  • Experience working in multi-timezone environments and coordinating across distributed teams.
  • Background establishing reliability practices and SRE frameworks in growing organizations.
Why Join Us:
  • 100% remote work with flexible hours
  • High-impact role with autonomy and ownership
  • Collaborative and international engineering team
  • Cutting-edge tech stack with strong focus on reliability and automation.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Platform Engineer (Kubernetes/CI-CD) - Hybrid Porto (3 days/week office)
Platform Engineer (Kubernetes/CI-CD) - Hybrid Porto (3 days/week office)

HumanIT Digital Consulting • Porto

Híbrido
EUR 52 000 - 78 000
Hybrid work model
Office in Porto
Senior SRE: Remote Cloud Reliability & Automation Lead
Senior SRE: Remote Cloud Reliability & Automation Lead

Pragmatike • Lisboa

Presencial
EUR 70 000 - 110 000
Remote work
Flexible hours
Autonomy
+2
Platform Engineer
Platform Engineer

Hexa Consulting • Portugal

Híbrido
EUR 50 000 - 70 000
Continuous learning and growth opportunities
Collaborative and innovation-driven environment
Real ownership and impact on platform evolution
Senior DevOps / SRE (Platform Reliability Engineer) - French fluent
Senior DevOps / SRE (Platform Reliability Engineer) - French fluent

emagine • Portugal

Presencial
EUR 60 000 - 90 000
Cloud Engineer (AWS/Kubernetes) - Full Remote Portugal
Cloud Engineer (AWS/Kubernetes) - Full Remote Portugal

HumanIT Digital Consulting • Porto

Teletrabalho
EUR 27 900 - 34 596
15th month salary
Health insurance (family)
Birthday off
+2
Cloudops Engineer
Cloudops Engineer

Sgi • Setúbal

Presencial
EUR 55 000 - 85 000
Site Reliability Engineer
Site Reliability Engineer

La Redoute • Leiria

Presencial
EUR 55 000 - 75 000
Systems Engineering, Metrics and Alerting
Systems Engineering, Metrics and Alerting

Cloudflare • Lisboa

Presencial
EUR 66 000 - 91 000
Equity plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Claranet Portugal • Portugal

Presencial
EUR 45 000 - 65 000
Senior DevOps Engineer – Cloud Platform
Senior DevOps Engineer – Cloud Platform

Solvace • Portugal

Teletrabalho
EUR 50 000 - 70 000