Senior Site Reliability Engineer

Jobgether

Portugal

Presencial

EUR 46 000 - 103 000

Tempo integral

Há 6 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Vantagens oferecidas por esta oferta de emprego

100% remote work
Flexible hours (async work)
Flexible paid time off
16 weeks parental leave
Home office budget and IT equipment
Stock options

Resumo da oferta

Jobgether is seeking a Senior Site Reliability Engineer based in Portugal for a fully remote role. You will own high-impact reliability projects spanning Kubernetes, AWS, IaC, observability, CI/CD, and AI-enabled workflows, with a global asynchronous team.

Mentor engineers and shape platform strategy while improving reliability across engineering teams. The role emphasizes SLOs/SLIs, incident response, and scalable cloud infrastructure, with autonomy and opportunities to influence architectural

Qualificações

  • Solid experience in Site Reliability Engineering, DevOps, or platform engineering.
  • Hands-on experience operating and scaling Kubernetes in production.
  • Proven production cloud infrastructure design and management using AWS or comparable provider.
  • Strong IaC skills with Terraform and infrastructure-as-code principles.
  • Experience with reliability frameworks (SLOs/SLIs, error budgets, alerting, incident management).
  • Strong observability with OpenTelemetry, Grafana, Prometheus, or similar.
  • Experience designing CI/CD pipelines (GitLab CI, GitHub Actions, etc.).
  • Programming in Golang, Bash, Python; scripting skills encouraged.

Responsabilidades

  • Lead discovery, design, and delivery of complex reliability initiatives.
  • Contribute to platform architecture, tooling, and engineering roadmaps.
  • Define and operate reliability practices including SLOs, SLIs, and alerting.
  • Use incident metrics to identify systemic reliability issues and guide strategy.
  • Automate recurring problems into reusable solutions, docs, and runbooks.
  • Operate and scale production Kubernetes environments and cloud infra.
  • Develop infrastructure as code using Terraform and support automated deployments.
  • Design AI-assisted infrastructure workflows and foster productivity.
  • Mentor less-senior engineers and collaborate with Security on hardening.
  • Contribute to capacity planning, performance optimization, and cost efficiency.
  • Participate in incident response and on-call rotations.

Conhecimentos

Kubernetes
Terraform
OpenTelemetry
Prometheus
CI/CD
Golang
Python
AWS

Formação académica

Bachelor's degree in CS or related field

Ferramentas

GitHub Actions
GitLab CI
Docker

Descrição da oferta de emprego

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Portugal.

This is a senior-level SRE opportunity focused on solving complex reliability, infrastructure, and platform challenges in a fully remote environment. You will take ownership of high-impact projects, from solution discovery through delivery, while helping shape platform architecture and long-term reliability strategy. Your work will span Kubernetes, cloud infrastructure, infrastructure as code, observability, CI/CD, and operational excellence. You will also play a key role in establishing strong SLOs, alerting, incident response, and reliability practices across engineering. AI is embedded into the way the team works, with an emphasis on building practical, reusable AI workflows that improve engineering productivity and safety. The role offers significant autonomy, technical influence, and opportunities to mentor engineers while collaborating asynchronously across a global organization.

Accountabilities
  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust and maintainable technical solutions.
  • Contribute to platform architecture, infrastructure tooling, technical roadmaps, and engineering priorities, advocating for initiatives that improve reliability and developer experience.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Use operational and incident metrics to identify systemic reliability issues and influence the team's technical strategy.
  • Resolve cross-team infrastructure and platform requests while turning recurring problems into reusable solutions, automation, documentation, and runbooks.
  • Operate and scale production Kubernetes environments and associated container infrastructure.
  • Build and manage cloud infrastructure using AWS or comparable cloud platforms, with strong emphasis on reliability, scalability, and operational efficiency.
  • Develop and maintain infrastructure as code using Terraform and support automated deployment workflows through modern CI/CD platforms.
  • Use AI natively in infrastructure, operations, and development workflows, creating reusable prompts, skills, tooling, and agentic workflows that improve team-wide productivity and reliability.
  • Design infrastructure and systems with AI-assisted engineering in mind, including clean interfaces, strong observability, secure-by-default patterns, CI protections, and review guardrails.
  • Mentor less-senior engineers through actionable feedback, technical guidance, hiring, onboarding, and RFC discussions.
  • Collaborate with Security on infrastructure hardening, threat mitigation, and defensive security practices.
  • Contribute to infrastructure capacity planning, performance optimization, and cost efficiency.
  • Participate in incident response and on-call rotations, helping maintain high standards of platform availability and reliability.
Requirements
  • Solid professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or a closely related discipline.
  • Strong hands-on experience operating and scaling Kubernetes in production, including Docker and the wider container ecosystem.
  • Proven experience designing, building, and managing production cloud infrastructure using AWS or a comparable cloud provider.
  • Strong practical expertise with Terraform and infrastructure-as-code principles.
  • Hands-on experience with reliability engineering frameworks, including SLOs, SLIs, error budgets, alerting, and incident management.
  • Strong observability experience with technologies such as OpenTelemetry, Grafana, Prometheus, or equivalent platforms.
  • Experience designing and operating CI/CD pipelines using GitLab CI, GitHub Actions, or similar technologies.
  • Comfortable working with Golang, Bash, and scripting, with broader programming experience considered an advantage.
  • Demonstrated practical use of AI and agentic workflows in infrastructure, operations, or software engineering, with measurable outcomes beyond simple familiarity with AI tools.
  • Strong understanding of production operations, troubleshooting, automation, and platform reliability.
  • Clear, thoughtful communication skills, particularly in an asynchronous and globally distributed environment.
  • Proactive and curious mindset, with the ability to independently identify problems, take ownership, and drive solutions to completion.
  • Collaborative and respectful approach when working across cultures, time zones, and diverse teams.
  • Experience with a backend programming language such as Elixir, Node.js, Python, or similar is a plus.
  • Experience operating and configuring Linux systems outside cloud environments is beneficial.
  • Security knowledge, including defensive and offensive security concepts, is an advantage.
Benefits
  • 100% remote work, with the ability to work from anywhere.
  • Flexible working hours within an async-first working environment.
  • Flexible paid time off to support a healthy balance between work and personal life.
  • 16 weeks of paid parental leave.
  • Budget for co-working spaces, learning, and wellness, including gym memberships.
  • Mental health support services.
  • Stock options.
  • Home office budget and IT equipment.
  • Competitive, location-aware compensation designed to support fair and equitable pay across global markets.
  • Annual salary range of $53,300–$119,850 USD, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.
  • Significant autonomy and ownership in a globally distributed engineering organization.
  • Opportunities to influence platform architecture, technical strategy, engineering standards, and reliability practices.
  • Supportive environment focused on professional growth, inclusion, collaboration, and continuous improvement.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Staff Application Engineer, Workplace Technology
Staff Application Engineer, Workplace Technology

Jobgether • Portugal

Teletrabalho
EUR 90 000 - 150 000
Generous performance-based bonus plans
Comprehensive medical coverage
Home office stipend
+5
Senior Software Engineer fullstack (Remote - Europe)
Senior Software Engineer fullstack (Remote - Europe)

Jobgether • Portugal

Teletrabalho
EUR 70 000 - 100 000
Competitive salary with performance-based bonuses
Flexible hybrid working environment
Comprehensive health insurance
+3
Site Reliability Engineering Manager (Data Infra)
Site Reliability Engineering Manager (Data Infra)

Complyadvantage • Lisboa

Híbrido
EUR 86 000 - 96 000
Equity participation
Private medical insurance
Unlimited Time Off Policy
+2
Senior Site Reliability Engineer - Remote AI-Driven Infra
Senior Site Reliability Engineer - Remote AI-Driven Infra

Jobgether • Portugal

Teletrabalho
EUR 46 000 - 103 000
100% remote work
Flexible hours (async work)
Flexible paid time off
+3
Site Reliability Engineer (SRE) @Lisboa
Site Reliability Engineer (SRE) @Lisboa

KCS iT • Lisboa

Híbrido
EUR 40 000 - 60 000
Programmes de formation gratuits
Expérience internationale
Options de travail flexible (hybride, à distance, sur site)
+1
Site Reliability Engineer
Site Reliability Engineer

TEKEVER • Lisboa

Presencial
EUR 60 000 - 90 000
Excellent work environment
Flexible work arrangements
Professional development opportunities
+1
Global IT Site Reliability Engineer Senior Manager
Global IT Site Reliability Engineer Senior Manager

Boston Consulting Group (BCG) • Lisboa

Presencial
EUR 60 000 - 80 000
AI Principal Engineer / Architect (EU Remote)
AI Principal Engineer / Architect (EU Remote)

Jobgether • Portugal

Teletrabalho
EUR 80 000 - 100 000
Competitive salary with performance-based opportunities
Flexible working arrangements
Professional development support
+2
Data and AI Engineer (EMEA)
Data and AI Engineer (EMEA)

Jobgether • Portugal

Teletrabalho
EUR 60 000 - 80 000
Competitive salary
Remote work flexibility
Professional development opportunities
Senior DevOps Engineer
Senior DevOps Engineer

Expert Executive Recruiters (EER Global) • Lisboa

Híbrido
EUR 65 000 - 90 000
Competitive compensation
Hybrid work model
Collaborative environment
+2