Senior Site Reliability Engineer

vistairhr

Barcelona

Híbrido

EUR 100.000 - 120.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Equipment: Laptop
Annual learning budget
Conference budget
AI tooling budget
Annual team offsite

Descripción de la vacante

Comply365 is seeking a Senior Site Reliability Engineer to strengthen our AWS-based platform, improve reliability, and mentor the infrastructure team. The role is based in Barcelona with a hybrid 2-office week schedule and core CET hours. You will architect, implement IaC, and collaborate with software engineers to increase production readiness.

You have 8+ years in senior SRE/infra roles, deep AWS expertise, and a proactive, automation-first mindset.

Formación

  • 8+ years in senior SRE, platform, or infra role for SaaS.
  • Proactive, automation-first mindset; solves problems before they escalate.
  • Independent work style with strategic thinking in complex environments.
  • Deep AWS expertise for scalable, resilient cloud architectures.
  • Experience migrating legacy infra to cloud-native environments.
  • Strong Linux, networking, and production resilience knowledge.
  • Proficient with IaC tools (Terraform, Pulumi) and CI/CD automation.
  • Strong communication and mentoring skills.

Responsabilidades

  • Identify risks and bottlenecks; implement durable, automated solutions.
  • Own reliability and infrastructure challenges end-to-end.
  • Design and optimize AWS infrastructure for scalability, resilience, and cost.
  • Lead migration of services to modern cloud-native architectures.
  • Develop and maintain Infrastructure as Code to automate provisioning.
  • Improve CI/CD pipelines for safe, fast deployments.
  • Mentor infrastructure engineers and promote best practices.
  • Eliminate toil by automating recurring tasks and documenting gaps.
  • Build observability across metrics, logs, traces, and alerts.
  • Collaborate with engineering to improve production readiness and system architecture.

Conocimientos

AWS
Terraform
Pulumi
CI/CD
SRE
Linux
Networking
Observability

Herramientas

Terraform
Pulumi
CI/CD tooling

Descripción del empleo

About Comply365

This is an excellent opportunity to join Comply365, the market leader in operational content management, safety management and training management serving the global aviation, defense and rail industries with a stellar customer base of 550+ customers in 80+ countries worldwide and delivering an industry first, connected platform, powered by AI, across these mission critical domains.

The Comply365 platform is an industry-first interconnected, intelligent platform with best-in-class solutions across Operational Content Management, Safety, Risk & Quality Management, and Training Management. Powered by industry-specific AI, our unified platform connects data and insights across operations, safety, and training - eliminating complexity and driving measurable gains in efficiency and productivity for aviation, rail and defense organizations worldwide.

About the role

At SafetyManager365, we move quickly, build with intent, and focus on solving problems that genuinely matter to our customers and the aviation industry. Airlines generate huge volumes of safety reports, operational data and investigations. Our platform helps safety teams analyse that information, identify risks earlier and make better operational decisions. As a newly formed team within the wider Comply365 group, we bring the energy and mindset of a startup, with the backing and stability of an established enterprise. Our AI and data processing workloads already run entirely on AWS, and we are now extending that approach as we modernise the rest of the platform. We are looking for a Senior Site Reliability Engineer to play a key role in this journey - enhancing reliability, shaping our cloud architecture, reducing operational toil, and mentoring our infrastructure team as they adopt modern cloud and SRE practices. This role is for someone who is genuinely energised by moving things forward - who spots a manual process and instinctively asks how it could be automated, and who pushes for durable solutions ("one-and-done") rather than repeating the same fixes. We are at an early stage of our cloud maturity journey and need an individual who is comfortable dealing with ambiguity, not someone who reaches for a checklist.

This is a high impact role that sits at the intersection of reliability engineering, cloud infrastructure, and engineering enablement, and will be ideal for someone who wants to combine technical leadership with hands-on engineering. You will play a key role in defining how we build, operate, and scale the platform. You will partner closely with software engineers to improve production readiness, deployment safety, observability, and operational maturity across the platform, reducing future incidents through better system design, simplification, and stronger operational foundations. The role is full-time (40 hours per week) with core hours from 10am to 5pm CET to ensure strong team alignment and collaboration, while still allowing flexibility outside of those hours. It is hybrid with at least 2 days in our offices in Barcelona per week and reports into the technical leadership within the SafetyManager365 group.

Please note this role is Hybrid based in Barcelona.

Key Responsibilities
  • Proactively identify risks, weaknesses, and operational bottlenecks, introducing structured solutions to move the platform from reactive to reliable and strategic
  • Take end-to-end ownership of reliability and infrastructure challenges, driving them through to durable resolution by treating root causes, not symptoms, and keeping standards above short-term workarounds
  • Design, build, and continuously improve AWS infrastructure with a focus on scalability, resilience, performance, and cost efficiency
  • Lead the migration of services and workloads from legacy dedicated environments into modern, cloud-native AWS architectures
  • Develop and maintain Infrastructure as Code to automate provisioning, reduce manual effort, and ensure consistency across environments. Proficiency in at least one scripting or general-purpose language (Python preferred) is required; you will be asked to work through a practical coding or automation problem as part of the interview process
  • Enhance CI/CD pipelines to enable safe, fast, and repeatable deployments, improving overall developer productivity and system stability
  • Mentor and support infrastructure engineers, promoting modern cloud and SRE practices, and leveraging AI tools where appropriate to improve automation, efficiency, and engineering outcomes
  • Identify and eliminate toil: where a task is recurring and automatable, automate it. Apply judgment to prioritise where automation delivers the most leverage, while documenting areas that aren’t yet, thereby avoiding single-points of failure.
  • Build and evolve observability capabilities across metrics, logging, tracing, and alerting to enable effective monitoring and rapid incident response
  • Establish best practices for incident management and collaborate with engineering teams to improve production readiness, resilience, performance, and system architecture
Skills and Qualifications
  • At least 8 years of experience in this or similar roles. Significant experience running production SaaS platforms in a senior SRE, platform, or infrastructure engineering role is required.
  • A genuinely proactive mindset - not just responsive, but anticipatory. You notice problems before they elevate, propose concrete solutions without being asked, and stay uncomfortable with the status quo. If you are content to do repetitive manual work rather than drive to automate it, this role is likely not for you.
  • Proven ability to operate independently and think strategically in complex, ambiguous environments, using sound judgement and strong problem-solving skills, without waiting for the whole picture
  • Deep, hands-on expertise with AWS, including designing and operating scalable, resilient, and cost-efficient cloud architectures
  • Demonstrated experience migrating and modernising legacy infrastructure into cloud-native environments
  • Strong foundations in Linux, networking, and systems internals, with experience designing for high availability, fault tolerance, and production resilience at scale.
  • Proficiency with Infrastructure as Code tools such as Terraform, Pulumi, or similar, alongside experience improving CI/CD and automation practices
  • Strong communication and collaboration skills, with a mentoring mindset coupled with experience of working closely with engineering teams to improve reliability, observability, and operational maturity.
Nice to haves
  • Deep expertise in AWS, with a proven ability to architect, operate, and optimise secure, scalable, resilient, and cost-efficient cloud platforms; AWS certifications are advantageous.
  • Proven track record defining and managing SLIs, SLOs, and error budgets, supported by hands-on experience with modern observability platforms (e.g., Prometheus, Grafana, Datadog) and distributed tracing.
  • Experience operating distributed systems and microservices architectures, including approaches such as service meshes, chaos engineering, and resilience testing.
  • Exposure to AI/ML or large-scale data workloads in production, with experience working in regulated or compliance-sensitive environments and a pragmatic approach to leveraging AI to improve engineering outcomes.
Benefits
  • Equipment: Laptop (Macbook), monitor, and whatever else you need to get productive
  • Annual learning budget: courses, books, conferences, coaching
  • Conference and speaking budget: attend industry events
  • AI tooling budget: pick the tools that make you better, we will cover them
  • Annual team offsite

Salary: €100,000 - €120,000 gross per annum. The salary offered will be determined within this range based on relevant experience, qualifications and demonstrated competencies.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Comply365 • Barcelona

Presencial
EUR 100.000 - 120.000
Laptop & monitor
Annual learning budget
Conference budget
+2
Senior Full Stack Engineer - EU hybrid
Senior Full Stack Engineer - EU hybrid

vistairhr • Barcelona

Híbrido
EUR 80.000 - 100.000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Comply365 • Barcelona

Híbrido
EUR 110.000 - 130.000
High-end equipment
Competitive salary
Annual offsite
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

vistairhr • Barcelona

Híbrido
EUR 110.000 - 130.000
High-end equipment
Competitive salary
Annual team offsite
+1
SRE/DevOps Engineer in Barcelona
SRE/DevOps Engineer in Barcelona

Blu Selection • Barcelona

Híbrido
EUR 60.000 - 75.000
Hybrid work model
Senior Sre Engineer - Tarragona
Senior Sre Engineer - Tarragona

Dempo • Tarragona

Híbrido
EUR 90.000 - 150.000
Private health insurance
Remote-friendly culture
Learning budget
+4
Senior Infrastructure Engineer
Senior Infrastructure Engineer

SITA • Barcelona

Presencial
EUR 60.000 - 80.000
Flex Week: Work from home up to 2 days/week
Flex Day: Make your workday suit your life
Flex-Location: Work from any location for 30 days a year
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tamarind Intelligence • Barcelona

Híbrido
EUR 55.000 - 70.000
Hybrid work model
Salary 55k-70k€ annually
Comprehensive health insurance
Senior SRE Engineer
Senior SRE Engineer

Dempo • España

Presencial
PHP 6.531.000 - 10.160.000
Remote-first culture
Granada office option
Private health insurance
+4
Senior Cloud SRE | AWS, AI-Driven Reliability | Hybrid
Senior Cloud SRE | AWS, AI-Driven Reliability | Hybrid

Comply365 • Barcelona

Híbrido
EUR 100.000 - 120.000
Laptop & monitor
Annual learning budget
Conference budget
+2