We are looking for a highly experienced Senior Infrastructure & Storage Engineer with strong hands-on expertise in Linux, on-premises infrastructure, distributed storage (Ceph/MinIO), and data-centre engineering. The ideal candidate should have a proven track record of designing, implementing, troubleshooting, and operationalizing mission-critical infrastructure environments.
Mandatory Skills & Experience
- 8+ years of experience in Infrastructure Engineering, Linux Administration, or Storage Engineering roles.
- Strong hands-on experience managing production on-premises data-centre environments.
- Expert-level administration and troubleshooting of RHEL/CentOS/Linux operating systems.
- Extensive experience with Ceph Storage, including:
- OSDs, MONs, Placement Groups
- Cluster health monitoring
- Replication and failure domains
- Capacity planning and performance tuning
- Recovery from degraded cluster states
- Experience managing MinIO distributed object storage environments.
- Strong understanding of:
- Storage architecture
- Backup & recovery strategies
- Disaster recovery planning
- Data protection and high availability
- Experience with physical server infrastructure:
- Bare-metal servers
- Storage controllers/HBAs
- RAID and disks
- Firmware upgrades
- Hardware lifecycle management
- Experience supporting virtualized infrastructure environments.
- Strong troubleshooting and RCA (Root Cause Analysis) skills for complex production incidents.
Kubernetes & Platform Infrastructure
- Experience supporting bare-metal Kubernetes environments.
- Knowledge of:
- CSI Drivers
- Persistent Volumes (PV)
- Persistent Volume Claims (PVC)
- StorageClasses
- Ceph-backed Kubernetes storage
- Ability to troubleshoot issues across infrastructure, storage, operating system, and Kubernetes layers.
Automation & Infrastructure as Code
- Hands-on experience with:
- Ansible
- Terraform
- Bash Shell Scripting
- Python Automation
- Experience using Git-based workflows for infrastructure management.
- Ability to automate provisioning, configuration, monitoring, and operational tasks.
Monitoring & Observability
- Experience with:Prometheus
- Grafana
- Elasticsearch
- New Relic
Strong knowledge of infrastructure monitoring, alerting, logging, capacity planning, and performance management.
Preferred/Good-to-Have Skills
- Rancher, RKE/RKE2 administration.
- MariaDB or PostgreSQL administration.
- Hybrid Azure infrastructure exposure.
- VMware or equivalent virtualization technologies.
- Windows Server administration (basic to intermediate).
- Multi-data-centre Disaster Recovery and Business Continuity exercises.