Senior Infrastructure & Storage Engineer

PSD Group

Barcelona

Presencial

EUR 60.000 - 90.000

Jornada completa

Hace 9 días
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

PSD Group is seeking a Senior Infrastructure & Storage Engineer to own the implementation, integration, testing and engineering lifecycle of on-premises compute, Linux, distributed storage and data-centre infrastructure in Barcelona. You will lead deployments, maintenance and handover with Platform Engineering and Security, focusing on Ceph, MinIO, and Kubernetes storage integration.

The role requires hands-on experience across bare-metal infrastructure, capacity planning, backup/recovery and

Formación

  • Strong hands-on experience engineering and troubleshooting production on-premises infrastructure.
  • Advanced Linux systems administration and troubleshooting capability.
  • Storage systems and data-protection concepts: replication, quorum, capacity, performance, backup and recovery.
  • Practical production experience with Ceph or similar distributed storage.
  • Experience with physical servers, disks/controllers, firmware, virtualization and data-center dependencies.
  • Automation of infrastructure via Ansible, Terraform, Bash, Python or equivalent.
  • Experience implementing monitoring, alerting, capacity management, and operational procedures.
  • Strong incident troubleshooting, root-cause analysis, risk management and vendor escalation skills.
  • Clear documentation and transfer of specialist knowledge.
  • Ceph administration in production, including recovery from degraded states and performance investigations.
  • Distributed MinIO on bare-metal infrastructure.
  • Kubernetes, CSI, StorageClasses and persistent volumes.
  • Prometheus, New Relic, Elasticsearch observability platforms.
  • Database administration: MariaDB, PostgreSQL.

Responsabilidades

  • Build, configure, troubleshoot, and lifecycle-manage physical and virtual servers across heterogeneous data-centre environments.
  • Engineer and support CentOS, RHEL, or equivalent Linux platforms including OS, services, packages, disks and performance.
  • Execute data-centre refreshes, hardware deployments, firmware and OS changes per approved designs.
  • Validate compute, network, storage, power, capacity, resilience, monitoring and dependencies before production acceptance.
  • Support Windows infrastructure where required.
  • Administer and troubleshoot Ceph, including cluster health, OSDs and monitors.
  • Engineer and support distributed MinIO deployments, object-storage availability and recovery.
  • Diagnose storage latency, degraded redundancy, disk/node failures and capacity risks.
  • Plan, document and test backup, DR procedures; distinguish platform redundancy from backup.
  • Coordinate escalation for high-risk or complex storage problems.
  • Support bare-metal Kubernetes nodes and their OS, hardware, network, container runtime and storage dependencies.
  • Engineer and troubleshoot Kubernetes persistent storage, CSI drivers, StorageClasses and persistent volumes.
  • Collaborate during cluster upgrades, node maintenance, capacity changes and stateful workload incidents.
  • Understand core Kubernetes concepts to diagnose failures across workload/node/network/CSI/storage.

Conocimientos

Linux administration
Ceph storage
MinIO storage
Kubernetes storage integration
Automation: Ansible, Terraform, Bash,/
Monitoring/observability
Hardware & virtualization
Root-cause analysis
Documentation/knowledge transfer

Herramientas

Ceph
MinIO
Kubernetes
Rancher
RKE2
CSI
Prometheus
New Relic
Elasticsearch
MariaDB/PostgreSQL

Descripción del empleo

Summary

Rate: Negotiable

Duration: 6 Months

About the Client

Our client is a global leader in aviation technology, delivering innovative IT and communications solutions that connect airlines, airports, aircraft and governments worldwide. Operating in over 200 countries and territories, they play a critical role in enabling the safe, secure and efficient movement of millions of passengers every day.

About the Role

The Senior Infrastructure & Storage Engineer owns the implementation, integration, testing, and engineering lifecycle of on-premises compute, Linux, distributed storage, and related data-centre infrastructure.

The role translates approved architecture into reliable, resilient, supportable platforms and provides deep engineering support for complex infrastructure incidents.

This is a senior, hands-on role with particular emphasis on bare-metal infrastructure, Ceph, MinIO, Linux, Kubernetes storage integration, capacity, backup and recovery, and operational knowledge transfer. The engineer works closely with Platform Architecture, Operations, Kubernetes & Platform Engineering, Network Engineering, Security, vendors, and delivery teams

Key Responsibilities
  • Build, configure, troubleshoot, and lifecycle-manage physical and virtual servers across heterogeneous data-centre environments.
  • Engineer and support CentOS, RHEL, or equivalent Linux platforms, including operating-system, service, package, certificate, disk, memory, process, and performance issues.
  • Execute data-centre refreshes, hardware deployments, firmware and operating-system changes, and replacement activities in accordance with approved designs.
  • Validate compute, network, storage, power, capacity, resilience, monitoring, and support dependencies before production acceptance.
  • Support selected Windows infrastructure where required.
Distributed Storage and Data Protection
  • Administer and troubleshoot Ceph, including cluster health, OSDs, monitors, placement groups, pools, capacity, performance, replication, and failure domains.
  • Engineer and support distributed MinIO deployments, object-storage availability, capacity healing, replication, and recovery.
  • Diagnose storage latency, degraded redundancy, disk and node failures, data-path issues, and capacity risks using an evidence-based approach.
  • Plan, document, and test backup, restoration, disaster recovery, and data-protection procedures; distinguish platform redundancy from backup.
  • Coordinate specialist and vendor escalation for high-risk or complex storage problems
  • Support bare-metal Kubernetes nodes and their operating-system, hardware, network, containerruntime, and storage dependencies.
  • Engineer and troubleshoot Kubernetes persistent storage, CSI drivers, StorageClasses, persistent volumes, mounts, and Ceph-backed workloads.
  • Collaborate with the Kubernetes & Platform Engineer during cluster upgrades, node maintenance, capacity changes, recovery testing, and stateful workload incidents.
  • Understand core Kubernetes concepts sufficiently to diagnose whether failures originate in the workload, node, network, CSI, storage, or underlying infrastructure.
Automation, Monitoring, and Capacity
  • Automate repeatable infrastructure provisioning and configuration using Terraform, Ansible, Bash, Python, or equivalent tools.
  • Use Git-based workflows so infrastructure changes are reviewable, auditable, and reproducible.
  • Implement Prometheus metrics, dashboards, and actionable alerts for hosts, hardware, Ceph, MinIO, storage paths, capacity, and backup health.
  • Track utilization, saturation, growth, hardware health, and redundancy; implement shortand mediumterm capacity actions.
  • Integrate infrastructure logs and telemetry with platforms such as New Relic and Elasticsearch where appropriate.
Testing, Handover, and Engineering Support
  • Plan and execute functional, integration, resilience, failover, capacity, backup, restoration, upgrade, and rollback testing.
  • Define acceptance criteria, evidence, abort conditions, maintenance procedures, and safe recovery paths.
  • Provide Level 3 support, lead root-cause analysis, and implement permanent corrective actions for complex infrastructure incidents.
  • Produce practical SOPs, runbooks, diagrams, asset and dependency records, maintenance procedures, and escalation paths.
  • Conduct hands-on knowledge transfer and ensure routine operations can be performed safely without dependency on one engineer.
What we are looking for
Required Skills

Strong recent, hands-on experience engineering and troubleshooting production on-premises infrastructure.

  • Advanced Linux systems administration and troubleshooting capability.
  • Strong experience with storage systems and data-protection concepts, including replication, quorum, failure domains, capacity, performance, backup, and recovery.
  • Practical production experience with Ceph or a closely comparable distributed storage platform.
  • Experience working with physical servers, disks, controllers/HBAs, firmware, operating systems, virtualization, and data-centre dependencies.
  • Experience automating infrastructure through Ansible, Terraform, Bash, Python, or equivalent technologies.
  • Experience implementing monitoring, alerting, capacity management, and operational procedures.
  • Strong incident troubleshooting, root-cause analysis, risk management, and vendor escalation skills.
  • Ability to produce clear documentation and transfer specialist knowledge effectively.
  • Direct administration of Ceph in production, including recovery from degraded states and performance investigations.
  • Distributed MinIO on dedicated servers or bare-metal infrastructure.
  • Kubernetes, Rancher, RKE2/RKE, CSI, and persistent-volume integration.
  • Prometheus, New Relic, Elasticsearch, or comparable observability platforms.
  • MariaDB, PostgreSQL, or other infrastructure database administration.
  • CentOS/RHEL, virtualization, selected Windows infrastructure, and hybrid Azure integration.
  • Business-continuity and disaster-recovery exercises across multiple data centres.

Candidates do not need to be application-platform or cloud specialists. Depth in Linux, on-premises infrastructure, distributed storage, safe recovery, and operational knowledge transfer is more important than matching every desirable product.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior On-Prem Infrastructure & Storage Engineer
Senior On-Prem Infrastructure & Storage Engineer

PSD Group • Barcelona

Presencial
EUR 60.000 - 90.000
Infrastructure Engineer
Infrastructure Engineer

Huxley • Barcelona

Híbrido
EUR 45.000 - 65.000
Infrastructure and Backup Engineer - Hybrid in Barcelona
Infrastructure and Backup Engineer - Hybrid in Barcelona

Michael Page • Barcelona

Presencial
EUR 80.000 - 110.000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • España

A distancia
EUR 70.000 - 100.000
Fully remote work environment
Technical leadership opportunities
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stellar Cyber • Banyoles

Presencial
EUR 90.000 - 130.000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+6
Senior Middle Level DevOps Engineer Full time - Hybrid Remote
Senior Middle Level DevOps Engineer Full time - Hybrid Remote

kentech-sp • España

Híbrido
EUR 55.000 - 80.000
Gym membership
Private health insurance
Pension plan
Sr. Staff Platform Operations Engineer
Sr. Staff Platform Operations Engineer

Cloudera • Madrid

Híbrido
EUR 90.000 - 125.000
Generous PTO Policy
Unplugged Days
Flexible WFH Policy
+6
Platform Engineer
Platform Engineer

Ilkari • Málaga

Presencial
EUR 55.000 - 75.000
Life and health insurance
Gym reimbursement
Learning Pocket budget
+4
Systems Engineer
Systems Engineer

Solera Holdings, LLC. • Madrid

Presencial
EUR 45.000 - 65.000
Site Reliability Engineer
Site Reliability Engineer

Emburse • Barcelona

Presencial
EUR 90.000 - 130.000
Flexible spending accounts
Generous paid time off
Paid parental leave
+9