Senior Site Reliability Engineer

outsystems

Portugal

Híbrido

EUR 60 000 - 90 000

Tempo integral

Há 3 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura feita para esta oferta — um currículo e uma carta de apresentação personalizados que vão ao encontro do anúncio.

Ultrapassa os filtros ATS

Resumo da oferta

OutSystems in Portugal is hiring a Site Reliability Engineer to strengthen the reliability of our production platforms. You will lead onboarding of services to reliability tenets and collaborate with software teams to ensure scalable, observable systems.

Hybrid / Onsite in Portugal (Lisbon, Braga, Proença-a-Nova). The role focuses on building cloud-native infrastructure, implementing SLOs/SLAs, monitoring, alerting, and on-call coverage, with Python and AI tooling to accelerate automation.

Qualificações

  • BS/MS in Computer Science or equivalent.
  • 6+ years of Site Reliability Engineering experience.
  • Experience managing Hadoop and Kubernetes infrastructure.
  • Advanced knowledge of Linux, Networking and Containers.
  • Proficiency in Python or Go.
  • Fluency in English and excellent communication skills.
  • Understanding or hands-on with Prompt engineering in software development.
  • Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude.

Responsabilidades

  • Lead and onboard services and teams to the reliability tenets.
  • Establish and maintain SLOs and SLAs.
  • Design and implement scalable, reliable cloud-native infrastructure.
  • Collaborate with software development teams to ensure resilient, observable systems.
  • Implement monitoring, alerting, logging, and tracing for incident detection.
  • Lead incident response and conduct RCA/post-mortems.
  • Automate operational tasks with focus on fast incident detection & recovery.
  • Program in Python with Gen AI tooling to accelerate automation.
  • Foster culture of continuous improvement and knowledge sharing.
  • Communicate reliability and performance updates to stakeholders.
  • Participate in on-call rotation for 24/7 production support.

Conhecimentos

SRE experience
Cloud infrastructure
Linux knowledge
Networking
Containers
Python/Go
English fluency
Prompt engineering
AI tooling awareness

Formação académica

BS/MS in Computer Science or Equivalent

Ferramentas

Kubernetes
Hadoop
IaC tools
EKS

Descrição da oferta de emprego

There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems!

Hybrid / Onsite in Portugal (Lisbon - Braga - Proença-a-Nova)

Site Reliability Engineer Role

As an SRE at OutSystems here are your key responsibilities and duties:

  • Lead and onboard services and teams to the reliability tenets;
  • Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs);
  • Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices;
  • Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant;
  • Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents;
  • Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems;
  • Automate every operational task, with a special focus on fast incident detection & recovery;
  • Programming in Python supported by Gen AI tooling to accelerate development of mission critical automation and tools.
  • Foster a culture of continuous improvement and knowledge sharing;
  • Communicate effectively with stakeholders, providing updates on system reliability and performance;
  • Participate in on-call rotation to provide 24/7 support for production systems.
Site Reliability Engineering Performance Indicators

The main KPIs that aid in understanding the impact and success of the SRE function at OutSystems are:

  • SLA and Service Level Objectives (SLO) compliance;
  • SLO Coverage and Detection Ratio;
  • MTTA - Mean time to acknowledge;
  • MTTR - Mean time to resolve.
Qualifications and Skills
Qualifications
  • BS/MS in Computer Science or Equivalent
  • 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale
  • History of end-to-end project delivery
  • Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience
  • Advanced knowledge of Linux, Networking, and Containers
  • Proficiency in at least one high-level programming language (Python, GoLang etc.).
  • Strong troubleshooting and debugging skills.
  • Fluency in English and excellent communication skills.
  • An understanding or hands‑on experience with Prompt engineering in software development;
  • Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude.
Soft Skills
  • Communication - able to communicate effectively (in English) both orally and written showing empathy for the other person;
  • Collaboration - Proactive collaboration and presentation skills to effectively communicate ideas and represent the deliverables and needs of the SRE team with leadership.
  • Humbleness - accepts mistakes and acts accordingly, with a humble attitude, apologizing for them and mitigating them ASAP to avoid higher impact.
  • Accountability - takes ownership of problems and makes sure to see them through. Even if he does not have all the necessary knowledge to move on alone, can involve the right people to reach closure.
  • Negotiation Skills - has tough and politically complex conversations with colleagues and customers, defusing disagreements and leading towards a mutual agreement and understanding of all parties involved.
  • Process Oriented - is organized and able to properly follow defined processes, whilst being able to properly challenge inefficient processes and suggest improvements.
  • Problem‑solving - Has a top‑down approach to problems, breaking them into smaller pieces and solving them by starting with a wider scope and narrowing it down as the analysis progresses. Has critical thinking, so can analyze information objectively and make a reasoned judgment.
Technical Skills
  • Experience in any of the following is valued, but not fully required: Ability to establish, monitor, and improve Service Level Objectives (SLOs), Indicators (SLIs), and Agreements (SLAs) in line with business needs.
  • Containerization technologies and orchestration platforms, mainly Kubernetes and EKS (CKA, CKAD, CKS certifications are valued);
  • Experience with automation and Infrastructure as Code (IaC) tools, such as A
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

OutSystems • Lisboa

Híbrido
EUR 70 000 - 110 000
Senior Site Reliability Engineer (Hybrid, Portugal)
Senior Site Reliability Engineer (Hybrid, Portugal)

OutSystems • Lisboa

Híbrido
EUR 70 000 - 110 000
Site Reliability Engineer (SRE) @Lisboa
Site Reliability Engineer (SRE) @Lisboa

KCS iT • Lisboa

Híbrido
EUR 40 000 - 60 000
Programmes de formation gratuits
Expérience internationale
Options de travail flexible (hybride, à distance, sur site)
+1
Site Reliability Engineer
Site Reliability Engineer

Fulcrum Digital Inc • Lisboa

Presencial
EUR 50 000 - 70 000
Site Reliability Engineer (Cloud & AI Platforms)
Site Reliability Engineer (Cloud & AI Platforms)

Komodo Consulting • Lisboa

Híbrido
EUR 50 000 - 75 000
Site Reliability Engineer
Site Reliability Engineer

La Redoute • Viseu

Presencial
EUR 60 000 - 86 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Claranet Portugal • Portugal

Presencial
EUR 45 000 - 65 000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobgether • Portugal

Teletrabalho
EUR 120 000 - 180 000
Fully remote
First dedicated SRE
Autonomy to shape practices
+1
Senior Site Reliability Engineer — Hybrid Porto
Senior Site Reliability Engineer — Hybrid Porto

Aubay Portugal • Porto

Híbrido
EUR 60 000 - 90 000
Health insurance
Training Academy
Career advancement
+2
SRE - Site Reliability Engineer - Kubernetes
SRE - Site Reliability Engineer - Kubernetes

Aubay Portugal • Porto

Híbrido
EUR 60 000 - 90 000
Health insurance
Training Academy
Career advancement
+2