Site Reliability Engineer

Valid

Madrid

Híbrido

EUR 70.000 - 100.000

Jornada completa

hace 26 horas
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Private medical insurance
Life insurance
Flexible hours

Descripción de la vacante

Valid is seeking a highly analytical Site Reliability Engineer to design and enhance the reliability architecture of our platforms. You will collaborate with R&D, DevOps, and Operations to ensure high availability, scalability, and observability across services.

The role requires 3+ years in SRE/DevOps, strong Linux and AWS skills, CI/CD experience, and a focus on performance and capacity planning. Spanish-based candidates with travel flexibility are preferred.

Formación

  • Bachelor’s degree in Computer Engineering, Electronics Engineering, Telecommunications Engineering, or related field.
  • 3+ years of experience in Site Reliability Engineering, Infrastructure Operations, DevOps, or similar.
  • 3+ years of experience within the telecommunications industry or related technology sectors.
  • Strong Linux administration skills (Red Hat, Ubuntu, or similar distributions).
  • Hands-on experience with cloud platforms, preferably AWS.
  • Experience designing and maintaining CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions, or similar).
  • Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack, Datadog, or equivalent).
  • Proficiency in scripting and automation using Python, Bash, or similar languages.
  • Experience with Infrastructure as Code (Terraform, Ansible, or equivalent).
  • Strong knowledge of containerization and orchestration technologies (Docker, Kubernetes).
  • Experience in performance monitoring, troubleshooting, and system optimization.
  • Knowledge of disaster recovery, backup strategies, and business continuity practices.
  • Experience working with SQL databases.
  • Advanced English communication skills (B2+/C1).
  • Candidates must be based in Spain or nearby European countries and be available to travel when required.

Responsabilidades

  • Site Reliability Engineering: develop, maintain reliable, scalable, and efficient systems with R&D.
  • Cross-functional Collaboration: improve CI/CD and automation with DevOps/Enablement/Operations teams.
  • Observability & Monitoring: implement metrics and monitor system performance.
  • Disaster Recovery & Backups: implement DR plans and robust backups.
  • Capacity Planning & Performance Engineering: forecast, scale, and optimize performance.
  • Migrations: analyze and plan complex migrations.

Conocimientos

Linux administration
AWS cloud
CI/CD pipelines
Monitoring/observability
Python/Bash scripting
Terraform/Ansible
Docker/Kubernetes
SQL databases
Advanced English

Educación

Bachelor’s degree in Computer/Electrical/Telecommunications Engineering

Herramientas

Docker
Kubernetes
Terraform
Ansible

Descripción del empleo

If you’re passionate about technology, innovative projects, and making a real impact, your place is here.

We are a global technology provider with 65+ years of experience, delivering a comprehensive portfolio of solutions across ID & Digital Government, Banking & Payments, and Trusted Connectivity. With more than 4,000 employees in 16 countries, we are committed to building a more secure and trustworthy world.

Within our Trusted Connectivity business unit, we develop cutting-edge solutions for the telecommunications industry—ranging from SIM cards and eSIMs to Subscription Management and secure connectivity services—connecting people, businesses, and devices worldwide.

We are looking for a highly analytical and business-oriented Site Reliability Engineer to design and enhance the reliability engineering architecture of our platforms, ensuring high availability, scalability, reliability, and observability through close collaboration with R&D, DevOps, and Operations teams.

As a Site Reliability Engineer, you will be responsible for working mainly with Site Reliability Architect (SRA) and R&D team to design resilient systems and operational processes that ensure the high availability, scalability, reliability and observability of our platforms.

What will you do?
Site Reliability Engineering
  • Work together with R&D to develop and maintain reliable, scalable, and efficient systems.
  • Work closely with R&D when new features are being developed and ensure that the new feature is ready to be released.
  • Ensure new features have been validate in terms of performance, reliability and saclability
  • Prepare and conduct knowledge transfer, documentation and information sharing to the other team members.
Cross-functional Collaboration
  • Work together with DevOps team to improve existing and implement new, effective CI/CD processes.
  • Work together with Enablement engineer to produce automation tools needed for performance and reliability monitoring
  • Work together with Operations team to support the platforms in terms of operational aspects.
  • Continuously evaluate and optimize system performance and capacity in order to maintain stable production platforms.
  • Identify, assess, and implement measures to eliminate potential risks that could impact the performance of systems and services.
  • Research, evaluate, test and advise at selecting appropriate new technologies or tools for improving site reliability
Observability & Monitoring
  • Monitor system performance, identifying bottlenecks, and execute pipeline optimization
  • Implement comprehensive service metrics to track and report on system reliability, performance, and efficiency.
Disaster Recovery & Backups
  • Implement disaster recovery plans and ensuring robust backup systems are in place.
Capacity Planning & Performance Engineering
  • Support in forecasting, scaling, and performance tuning.
  • Create KPI to monitor growth and optimize resource utilization.
Migrations
  • Analyze and plan for complex migrations.
What are we looking for?
  • Bachelor’s degree in Computer Engineering, Electronics Engineering, Telecommunications Engineering, or a related field.
  • 3+ years of experience in Site Reliability Engineering, Infrastructure Operations, DevOps, or a similar role.
  • 3+ years of experience within the telecommunications industry or related technology sectors.
  • Strong Linux administration skills (Red Hat, Ubuntu, or similar distributions).
  • Hands-on experience with cloud platforms, preferably AWS.
  • Experience designing and maintaining CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions, or similar).
  • Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack, Datadog, or equivalent).
  • Proficiency in scripting and automation using Python, Bash, or similar languages.
  • Experience with Infrastructure as Code (Terraform, Ansible, or equivalent).
  • Strong knowledge of containerization and orchestration technologies (Docker, Kubernetes).
  • Experience in performance monitoring, troubleshooting, and system optimization.
  • Knowledge of disaster recovery, backup strategies, and business continuity practices.
  • Experience working with SQL databases.
  • Advanced English communication skills (B2+/C1).
  • Candidates must be based in Spain or nearby European countries and be available to travel when required.
If you want this position to be yours, we would like you to have the following:
  • AWS, Linux, or Kubernetes certifications.
  • Experience in highly available and mission-critical environments.
  • Knowledge of capacity planning and performance engineering.
  • Experience in telecom platforms, mobile services, or cloud-native architectures.
What we offer
  • Join Valid and work on innovative, global technology projects within multicultural and multidisciplinary teams.
  • Flexibility: flexible working hours and remote work options to support work-life balance.
  • Well-being first: private medical insurance and life insurance.
  • Be part of a company that values continuous learning, collaboration, and growth.
Our Culture

At Valid, we foster an inclusive, diverse, and innovative environment where everyone thrives. We are committed to equal opportunities, free from discrimination concerning sex, age, race, sexual orientation, religion, education, social status, culture, or special needs such as illness or disability. We value people as the heart of our culture. Trust, transparency, and teamwork are the foundations of our success, driving growth and empowering talent.

Join this great team and be part of our story!

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Field Application Engineer
Field Application Engineer

Valid • Madrid

Híbrido
EUR 50.000 - 75.000
Hybrid model
Private medical insurance
Life insurance
+2
Site Reliability Engineer - DevSecOps Engineer
Site Reliability Engineer - DevSecOps Engineer

Kyndryl • España

Híbrido
EUR 65.000 - 90.000
Senior Site Reliability Engineer (d/f/m)
Senior Site Reliability Engineer (d/f/m)

TK Elevator • Madrid

Híbrido
EUR 60.000 - 80.000
Health and Safety activities
Flexible working hours
Training and education programs
+1
Jr. Quality Assurance Engineer
Jr. Quality Assurance Engineer

Valid • Madrid

Híbrido
EUR 30.000 - 45.000
Private medical insurance
Life insurance
Hybrid work model
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cabify • Madrid

Híbrido
EUR 45.000 - 75.000
Flexible hours
Team events
Staff free rides
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F. Hoffmann-La Roche AG • Sant Cugat del Vallès

Presencial
EUR 50.000 - 75.000
Site Reliability Engineer with Node.js
Site Reliability Engineer with Node.js

Kodify Media Group • Barcelona

A distancia
EUR 45.000 - 65.000
Fully remote position
Flexible working hours
10% on top of your salary for learning
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Comply365 • Barcelona

Presencial
EUR 100.000 - 120.000
Laptop & monitor
Annual learning budget
Conference budget
+2
Application Specialist
Application Specialist

Verisure • Alicante

Presencial
EUR 42.000 - 56.000
Lunch included
Dynamic environment
International projects
+2
Site Reliability Engineer - Data Platform
Site Reliability Engineer - Data Platform

N26 • Barcelona

Presencial
EUR 75.000 - 95.000
Premium bank account access
Work from home budget
Visa/relocation support
+1