Site Reliability Engineering Manager (Data Infra)

Complyadvantage

Lisboa

Presencial

EUR 86 000 - 96 000

Tempo integral

14 dias+
Gerador de candidaturas

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Equity participation
Private medical insurance
Unlimited Time Off Policy
Home office setup budget
Annual learning budget

Resumo da oferta

Complyadvantage in Lisbon is seeking a Site Reliability Engineering Manager to lead a team of SREs and shape the reliability strategy. You will work closely with various teams to ensure system resilience, scalability, and security.

This role involves managing engineering teams, improving system performance, and promoting best practices in observability and automation. Candidates should have strong experience with Kubernetes and Terraform, alongside excellent communication skills.

The position also offers a competitive salary range of €86,000 - €96,000, along with equity and additional benefits.

Qualificações

  • Experience managing and growing engineering teams.
  • Experience with Kubernetes and Terraform.
  • Experience in hosting microservices-based architectures.
  • Experience with cloud-native architectures (AWS and GCP preferred).
  • Good communication and writing skills.

Responsabilidades

  • Lead and empower a team of SREs.
  • Shape reliability strategy and improve system performance.
  • Mentor and coach your team while fostering a positive environment.
  • Set the direction for your team and take accountability for tech decisions.
  • Collaborate with stakeholders to meet internal customer needs.

Conhecimentos

Experience managing high performing engineering teams
Kubernetes
Terraform
Cloud native architectures (AWS, GCP)
Technical documentation writing

Ferramentas

GitOps
Grafana
ArgoCD

Descrição da oferta de emprego

We are looking for a driven and experienced Site Reliability Engineering Manager to join our innovative Tech Team. You will lead and empower a team of SREs, partnering closely with Engineering, Product, and Security to ensure our platforms are resilient, scalable, and secure. You will play a key role in shaping reliability strategy, improving system performance, and embedding best practices in observability, incident management, and automation to support the delivery of high-impact solutions in the fight against financial crime.

As a Site Reliability Engineering (SRE) Manager, you will:
  • Take ownership of your team, being responsible for current team members’ growth and development, plus hiring and onboarding new team members
  • Create a positive environment where your team members thrive to deliver the best outcomes and innovations
  • Be a role model for your team, mentoring and coaching them, whilst having a learning mindset yourself, being open to new ideas and technologies
  • Within the context of our broader technology vision, set the direction for your team and take accountability for tech decisions
  • Use your specific experience working with cloud systems to input into technical decision-making
  • Work with other stakeholders across engineering to ensure the systems and services your team provides meet the needs of your internal customers
  • Collaborate, both within your team and across the tribe to ensure your team’s implementation meets industry standards

The role reports to the Director of Infrastructure. You’ll be managing a team of Engineers focused on the provision and support of our Stateful / Data layer technologies powering all of our services, both in development and production. The main technologies we use are YugaByte (sharded Postgres), Kafka (via Strimzi), Elasticsearch (via ECK), Redis and Spark/data warehousing on GCP and AWS using their PaaS systems. As the technology stack underpins all other engineering work, a collaborative mindset is a must.

Our tech stack

ComplyAdvantage is fully cloud-based, with a modern kubernetes-focused tech stack. All compute workloads run in Kubernetes, with clusters in multiple regions to support the needs of our global client base. Our production services are multi-cloud by design and are currently hosted in AWS and GCP.

We make heavy use of Terraform and Helm to define our infrastructure and services, and lean heavily on GitOps paradigms - production and non-production environments are defined in git and changes to these environments (both cloud infrastructure and Kubernetes applications) are managed via git.

ArgoCD is our tool of choice for controlling our deployments, and paired with our Istio mesh, allows us for advanced deployment patterns used by our development teams such a progressive rollouts.

Our observability stack consists of Grafana Cloud, along with some on-prem Mimir, amongst others. We focus on Open Telemetry for application metrics, with SLO and metric driven alerting at all levels, from Cloud infra through to application performance.

Across the wider Technology team, teams build and release containerised applications to support the wide array of activities that our teams are engaged in - from developing low latency client-facing APIs, to machine learning models and data processing pipelines.

About you

As an Site Reliability Engineering (SRE) Manager, you will

  • Have experience of managing and growing high performing engineering teams
  • Have experience with Kubernetes and Terraform
  • Have experience hosting microservices-based architectures
  • Have experience of working with cloud native architectures (AWS and GCP are preferred)
  • Have good communication and writing skills including experience writing technical documentation
Nice to haves
  • Experience of working in a start-up/ scale-up environment
  • Have experience managing observability platforms, whether self-hosted or third party - eg Grafana stack, Datadog, NewRelic
  • Have experience managing pipeline tools, whether self-hosted or third party - eg CircleCI, ArgoCD, Harness, etc
Benefits
  • Equity as we want you to have a part of what we are building
  • Private medical insurance designed to keep you ensuring peace of mind while you excel in your career
  • Unlimited Time Off Policy- A work-life balance and focus on our well-being are critical to keeping us performing at our best
  • We embrace a hybrid approach that requires employees to be in the office for two days a week. We strongly believe that this approach fosters collaboration and enables the building of meaningful relationships
  • You will also get a new starter budget to kit out your home office
  • Opportunity to work on innovative projects with smart-minded people keen to share their knowledge and continuously improve
  • Annual learning budget (prorated based on start date) to drive your performance and career development

Base salary range for this role is86,000-96,000EUR+ equity and benefits. The actual pay may vary based on factors such as location, experience, and skills.

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Site Reliability Engineer - Platform Reliability (Resilience)
Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic • Portugal

Teletrabalho
EUR 63 000 - 84 000
Site Reliability Engineer
Site Reliability Engineer

TEKEVER • Lisboa

Presencial
EUR 60 000 - 90 000
Excellent work environment
Flexible work arrangements
Professional development opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Claranet Portugal • Portugal

Presencial
EUR 45 000 - 65 000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Devexperts • Porto

Presencial
EUR 55 000 - 75 000
Hybrid work mode
Work From Anywhere Program
Medical insurance for employees and 
2
+11
Global IT Site Reliability Engineer Senior Manager
Global IT Site Reliability Engineer Senior Manager

Boston Consulting Group (BCG) • Lisboa

Presencial
EUR 60 000 - 80 000
Systems Engineer – SRE
Systems Engineer – SRE

Matchtech • Lisboa

Presencial
EUR 55 000 - 65 000
Private healthcare plan
Meal allowance
Profit-sharing opportunities
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Devexperts LLC • Porto

Presencial
EUR 60 000 - 90 000
Hybrid work mode
Wellness days
Medical insurance
+5
Senior Site Reliability Engineer - PSRE
Senior Site Reliability Engineer - PSRE

Arcesium LLC • Lisboa

Presencial
EUR 50 000 - 80 000
Flexible work arrangements
Competitive compensation and benefits
Continuous learning and development opportunities
Observability Engineer - Site Reliability Engineer
Observability Engineer - Site Reliability Engineer

La Redoute • Viseu

Híbrido
EUR 60 000 - 100 000
Site Reliability Engineering
Site Reliability Engineering

Amadeus • Lisboa

Presencial
EUR 42 000 - 60 000
Global DNA
Learning opportunities
Flexible working model
+1