Senior Site Reliability Engineer

Comply365

Barcelona

Presencial

EUR 100.000 - 120.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Laptop & monitor
Annual learning budget
Conference budget
AI tooling budget
Annual team offsite

Descripción de la vacante

Comply365 in Barcelona is seeking a Senior Site Reliability Engineer to strengthen our cloud‑native platform and lead reliability initiatives. You will own AWS architecture, improve observability, and drive automation to reduce toil in a fast‑growing SaaS environment.

Hybrid role based in Barcelona, with mentoring responsibilities and collaboration across engineering teams to deliver durable solutions and scalable infrastructure.

Formación

  • 8+ years in senior SRE, platform, or infrastructure engineering roles.
  • Proactive and autonomous with strong problem-solving skills.
  • Ability to operate independently in complex, ambiguous environments.
  • Deep expertise with AWS cloud architecture and scalable design.
  • Experience migrating legacy infrastructure to cloud-native environments.
  • Solid Linux, networking, and systems internals knowledge.
  • Proficiency with IaC tools such as Terraform or Pulumi and CI/CD improvements.
  • Strong communication and collaboration for mentoring engineering teams.

Responsabilidades

  • Identify risks and bottlenecks, moving platform from reactive to reliable and strategic.
  • Own end-to-end reliability and infrastructure challenges; drive durable resolutions.
  • Design, build, and optimize AWS infrastructure for scalability, resilience, and cost efficiency.
  • Lead migration of services to cloud-native AWS architectures and modern environments.
  • Develop and maintain Infrastructure as Code to automate provisioning and reduce manual effort.
  • Enhance CI/CD pipelines for safe, fast, repeatable deployments.
  • Mentor infrastructure engineers and promote modern cloud and SRE practices.
  • Identify and eliminate toil through automation and documentation of non-automatable areas.
  • Develop observability across metrics, logging, tracing, and alerting for rapid incident response.
  • Establish incident management best practices with engineering teams.

Conocimientos

AWS
Infrastructure as Code
Terraform
Pulumi
CI/CD
Observability
Linux
Networking
Mentoring
Cost optimization

Herramientas

Prometheus
Grafana
Datadog
Kubernetes

Descripción del empleo

About Comply365

This is an excellent opportunity to join Comply365, the market leader in operational content management, safety management and training management serving the global aviation, defense and rail industries with a stellar customer base of 550+ customers in 80+ countries worldwide and delivering an industry first, connected platform, powered by AI, across these mission critical domains.

The Comply365 platform is an industry-first interconnected, intelligent platform with best-in-class solutions across Operational Content Management, Safety, Risk & Quality Management, and Training Management. Powered by industry-specific AI, our unified platform connects data and insights across operations, safety, and training – eliminating complexity and driving measurable gains in efficiency and productivity for aviation, rail and defense organizations worldwide.

About the role

At SafetyManager365, we move quickly, build with intent, and focus on solving problems that genuinely matter to our customers and the aviation industry. Airlines generate huge volumes of safety reports, operational data and investigations. Our platform helps safety teams analyse that information, identify risks earlier and make better operational decisions. As a newly formed team within the wider Comply365 group, we bring the energy and mindset of a startup, with the backing and stability of an established enterprise. Our AI and data processing workloads already run entirely on AWS, and we are now extending that approach as we modernise the rest of the platform. We are looking for a Senior Site Reliability Engineer to play a key role in this journey – enhancing reliability, shaping our cloud architecture, reducing operational toil, and mentoring our infrastructure team as they adopt modern cloud and SRE practices. This role is for someone who is genuinely energised by moving things forward – who spots a manual process and instinctively asks how it could be automated, and who pushes for durable solutions (“one-and-done”) rather than repeating the same fixes. We are at an early stage of our cloud maturity journey and need an individual who is comfortable dealing with ambiguity, not someone who reaches for a checklist.

Please note this role is Hybrid based in Barcelona.

Key Responsibilities
  • Proactively identify risks, weaknesses, and operational bottlenecks, introducing structured solutions to move the platform from reactive to reliable and strategic.
  • Take end-to-end ownership of reliability and infrastructure challenges, driving them through to durable resolution by treating root causes, not symptoms, and keeping standards above short‑term workarounds.
  • Design, build, and continuously improve AWS infrastructure with a focus on scalability, resilience, performance, and cost efficiency.
  • Lead the migration of services and workloads from legacy dedicated environments into modern, cloud‑native AWS architectures.
  • Develop and maintain Infrastructure as Code to automate provisioning, reduce manual effort, and ensure consistency across environments. Proficiency in at least one scripting or general‑purpose language (Python preferred) is required; you will be asked to work through a practical coding or automation problem as part of the interview process.
  • Enhance CI/CD pipelines to enable safe, fast, and repeatable deployments, improving overall developer productivity and system stability.
  • Mentor and support infrastructure engineers, promoting modern cloud and SRE practices, and leveraging AI tools where appropriate to improve automation, efficiency, and engineering outcomes.
  • Identify and eliminate toil: where a task is recurring and automatable, automate it. Apply judgment to prioritise where automation delivers the most leverage, while documenting areas that aren’t yet, thereby avoiding single‑points of failure.
  • Build and evolve observability capabilities across metrics, logging, tracing, and alerting to enable effective monitoring and rapid incident response.
  • Establish best practices for incident management and collaborate with engineering teams to improve production readiness, resilience, performance, and system architecture.
Skills and Qualifications
  • At least 8 years of experience in this or similar roles. Significant experience running production SaaS platforms in a senior SRE, platform, or infrastructure engineering role is required.
  • A genuinely proactive mindset – not just responsive, but anticipatory. You notice problems before they escalated, propose concrete solutions without being asked, and stay uncomfortable with the status quo. If you are content to do repetitive manual work rather than drive to automate it, this role is likely not for you.
  • Proven ability to operate independently and think strategically in complex, ambiguous environments, using sound judgement and strong problem‑solving skills, without waiting for the whole picture.
  • Deep, hands‑on expertise with AWS, including designing and operating scalable, resilient, and cost‑efficient cloud architectures.
  • Demonstrated experience migrating and modernising legacy infrastructure into cloud‑native environments.
  • Strong foundations in Linux, networking, and systems internals, with experience designing for high availability, fault tolerance, and production resilience at scale.
  • Proficiency with Infrastructure as Code tools such as Terraform, Pulumi, or similar, alongside experience improving CI/CD and automation practices.
  • Strong communication and collaboration skills, with a mentoring mindset coupled with experience working closely with engineering teams to improve reliability, observability, and operational maturity.
Nice to haves
  • Deep expertise in AWS, with a proven ability to architect, operate, and optimise secure, scalable, resilient, and cost‑efficient cloud platforms; AWS certifications are advantageous.
  • Proven track record defining and managing SLIs, SLOs, and error budgets, supported by hands‑on experience with modern observability platforms (e.g., Prometheus, Grafana, Datadog) and distributed tracing.
  • Experience operating distributed systems and microservices architectures, including approaches such as service meshes, chaos engineering, and resilience testing.
  • Exposure to AI/ML or large‑scale data workloads in production, with experience working in regulated or compliance‑sensitive environments and a pragmatic approach to leveraging AI to improve engineering outcomes.
Benefits
  • Equipment: Laptop (Macbook), monitor, and whatever else you need to get productive.
  • Annual learning budget: courses, books, conferences, coaching.
  • Conference and speaking budget: attend industry events.
  • AI tooling budget: pick the tools that make you better, we will cover them.
  • Annual team offsite.

Salary: €100,000 - €120,000 gross per annum.

The salary offered will be determined within this range based on relevant experience, qualifications and demonstrated competencies.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Full Stack Engineer - EU hybrid
Senior Full Stack Engineer - EU hybrid

Comply365 • Barcelona

Híbrido
EUR 80.000 - 100.000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

SITA • Barcelona

Presencial
EUR 60.000 - 80.000
Flex Week: Work from home up to 2 days/week
Flex Day: Make your workday suit your life
Flex-Location: Work from any location for 30 days a year
Senior Cloud SRE | AWS, AI-Driven Reliability | Hybrid
Senior Cloud SRE | AWS, AI-Driven Reliability | Hybrid

Comply365 • Barcelona

Híbrido
EUR 100.000 - 120.000
Laptop & monitor
Annual learning budget
Conference budget
+2
Principal CloudOps Engineer
Principal CloudOps Engineer

Sage • Barcelona

Híbrido
EUR 75.000 - 100.000
Medical and dental insurance
Flexible benefits
Well-being support
+4
Engineering Manager – Access & Control Platform
Engineering Manager – Access & Control Platform

Jobgether • Madrid

A distancia
EUR 110.000 - 140.000
Fully remote Europe
Flexible hours
Private health insurance
+7
Senior DevOps Engineer
Senior DevOps Engineer

Swiss Re • Madrid

Presencial
EUR 60.000 - 100.000
Variable compensation based on performance
Global and location-specific benefits
Senior DevOps Engineer
Senior DevOps Engineer

Swiss Re - Schweizerische Rückversicherungs-Gesellschaft • Madrid

Presencial
EUR 60.000 - 100.000
Performance-based variable compensation
Global and location-specific benefits
Site Reliability Engineer
Site Reliability Engineer

Valid • Madrid

Híbrido
EUR 70.000 - 100.000
Private medical insurance
Life insurance
Flexible hours
Site Reliability Engineer
Site Reliability Engineer

Restb • Bellprat

Híbrido
EUR 60.000 - 90.000
Hybrid work policy
Health insurance
Free in-office snacks, beverages and a
Cloud Security Engineer
Cloud Security Engineer

Montash • Barcelona

Híbrido
EUR 70.000 - 95.000
Health insurance
Learning budget
Flexible working