Senior DevOps / SRE (Platform Reliability Engineer) - French fluent

emagine

Portugal

Presencial

EUR 60 000 - 90 000

Tempo integral

há 12 horas
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura feita para esta oferta — um currículo e uma carta de apresentação personalizados que vão ao encontro do anúncio.

Ultrapassa os filtros ATS

Resumo da oferta

emagine is seeking a Senior DevOps / Site Reliability Engineer (SRE) to ensure reliability, scalability, performance, and security of our cloud platform. You will build and operate cloud-native systems, improve observability, automate operations, and support development teams for highly available services.

Key responsibilities include AWS infrastructure management, CI/CD, IaC with Terraform, Kubernetes orchestration, monitoring/logging, and on-call incident response.

Qualificações

  • 5+ years of experience in DevOps / SRE / Cloud Infrastructure / Platform Engineering.
  • Strong expertise in Linux systems administration and troubleshooting.
  • Proven experience with Kubernetes in production environments.
  • Strong experience with CI/CD tools (GitLab CI, Jenkins, GitHub Actions, Azure DevOps).
  • Solid knowledge of Infrastructure as Code (Terraform highly preferred).
  • Experience with AWS cloud platforms.
  • Strong understanding of networking fundamentals (TCP/IP, DNS, load balancing, reverse proxies).
  • Experience with observability tools: monitoring, metrics, logging, tracing.
  • Strong scripting skills (Bash, Python, or similar).
  • French advanced level.

Responsabilidades

  • Design, implement, and maintain highly available and scalable infrastructure on AWS.
  • Own and improve the reliability of production systems using SRE principles (SLO, SLI, error budgets).
  • Build and manage CI/CD pipelines to support fast and safe software delivery.
  • Develop and maintain Infrastructure as Code (IaC) using Terraform, Ansible, CloudFormation, etc.
  • Manage and optimize container orchestration platforms (Kubernetes, Docker, Helm).
  • Implement and maintain monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK, Datadog, Splunk).
  • Lead incident response, perform root cause analysis, and write postmortems to drive continuous improvement.
  • Improve system performance, capacity planning, scaling strategies, and disaster recovery processes.
  • Collaborate closely with development teams to improve deployment strategies and system resilience.
  • Implement security best practices (IAM, secret management, vulnerability scanning, patching).
  • Define operational standards, runbooks, documentation, and best practices for platform reliability.
  • Participate in on-call rotation and provide senior-level support for critical production issues.

Conhecimentos

Linux administration
SRE practices
Cloud infrastructure
Scripting (Bash/Python)
Incident management
On-call readiness

Ferramentas

Kubernetes
GitLab CI
Jenkins
GitHub Actions
Azure DevOps
Terraform
Ansible
CloudFormation
Prometheus
Grafana
ELK
Datadog
Splunk
Docker
Helm

Descrição da oferta de emprego

We are looking for a Senior DevOps / Site Reliability Engineer (SRE) to ensure the reliability, scalability, performance, and security of our platform and cloud infrastructure. You will play a key role in building and operating cloud-native systems, improving observability, automating operations, implementing SRE best practices (SLOs/SLIs), and supporting development teams to deliver highly available services.

Key Responsibilities
  • Design, implement, and maintain highly available and scalable infrastructure on AWS.
  • Own and improve the reliability of production systems using SRE principles (SLO, SLI, error budgets).
  • Build and manage CI/CD pipelines to support fast and safe software delivery.
  • Develop and maintain Infrastructure as Code (IaC) using Terraform, Ansible, CloudFormation, etc.
  • Manage and optimize container orchestration platforms (Kubernetes, Docker, Helm).
  • Implement and maintain monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK, Datadog, Splunk).
  • Lead incident response, perform root cause analysis, and write postmortems to drive continuous improvement.
  • Improve system performance, capacity planning, scaling strategies, and disaster recovery processes.
  • Collaborate closely with development teams to improve deployment strategies and system resilience.
  • Implement security best practices (IAM, secret management, vulnerability scanning, patching).
  • Define operational standards, runbooks, documentation, and best practices for platform reliability.
  • Participate in on-call rotation and provide senior-level support for critical production issues.
Key Responsibilities (5 Main Missions)
Mission 1: AWS Infrastructure Management (Build & Run)
Mission 2: CI/CD and Deployment Automation
Mission 3: Monitoring, Observability, and Alerting: Global Monitoring
Log Management
Application Monitoring
Business Analytics
Mission 4: Incident Management, Resilience, and Security
Mission 5: FinOps and AWS Cost Optimization
Key Requirements
  • 5+ years of experience in DevOps / SRE / Cloud Infrastructure / Platform Engineering.
  • Strong expertise in Linux systems administration and troubleshooting.
  • Proven experience with Kubernetes in production environments.
  • Strong experience with CI/CD tools (GitLab CI, Jenkins, GitHub Actions, Azure DevOps).
  • Solid knowledge of Infrastructure as Code (Terraform highly preferred).
  • Experience with AWS cloud platforms.
  • Strong understanding of networking fundamentals (TCP/IP, DNS, load balancing, reverse proxies).
  • Experience with observability tools: monitoring, metrics, logging, tracing.
  • Strong scripting skills (Bash, Python, or similar).
  • French advanced level.
Nice to Have
  • Experience with additional cloud platforms (Azure, GCP).
  • Strong understanding of networking fundamentals.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Site Reliability Engineer
Site Reliability Engineer

La Redoute • Leiria

Presencial
EUR 55 000 - 75 000
Site Reliability Engineer
Site Reliability Engineer

La Redoute • Viseu

Presencial
EUR 60 000 - 86 000
Platform Engineer
Platform Engineer

Damia • Lisboa

Presencial
EUR 70 000 - 100 000
DevOps Engineer - AWS / Kubernetes
DevOps Engineer - AWS / Kubernetes

act digital EMEA - Alter Solutions • Lisboa

Presencial
EUR 50 000 - 75 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Claranet Portugal • Portugal

Presencial
EUR 45 000 - 65 000
Senior DevOps Engineer – Cloud Platform
Senior DevOps Engineer – Cloud Platform

Solvace • Portugal

Teletrabalho
EUR 50 000 - 70 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Arcesium • Lisboa

Presencial
EUR 60 000 - 90 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

outsystems • Portugal

Híbrido
EUR 60 000 - 90 000
DevOps - Lisbon
DevOps - Lisbon

Inetum • Lisboa

Presencial
EUR 55 000 - 85 000
Devsecops Engineer (French Speaker)
Devsecops Engineer (French Speaker)

Inetum Portugal • Lisboa

Híbrido
EUR 60 000 - 90 000
Hybrid work model