Location Brussels-50% onsite
Contract 6-12 months contract (potential extensions)
About the role
We are looking for an Operations Manager to take over leadership of the team that runs our on-premise infrastructure: the Linux server estate, virtualisation, storage and the network elements that underpin mission-critical services. These services run 24/7 and downtime has real operational consequences.
The role has two halves. First, own and tighten day-to-day operations: incident, change and configuration control, ITSM discipline in ServiceNow, and a team that knows exactly what is running, where, and why. Second, be the operational anchor for the move of this landscape away from on-premise hosting towards VMware-based cloud platforms (Azure VMware Solution or Amazon Elastic VMware Service). The migration programme is led elsewhere; this role brings the knowledge of the estate into it, sets the conditions under which the target platform is accepted into operations, and then runs it — while keeping the current estate stable throughout.
This is a hands-on management role for someone who is meticulous by nature, treats production with respect, and considers "we don't know" an unacceptable answer about a mission-critical system.
The environment
- Linux server estate (enterprise distributions), VMware-based virtualisation, storage and backup
- Network elements: switching, routing, firewalls, load balancers, DNS/DHCP/NTP, network segmentation
- ITSM on ServiceNow: incident, problem, change, configuration (CMDB), knowledge
- Target state: hybrid, then cloud-hosted VMware (AVS / EVS) with on-premise decommissioning — delivered by a migration programme that this role supports and takes into operations
Key responsibilities
Run the platform
- Own end-to-end availability, performance and security posture of the on-premise Linux and network infrastructure against agreed SLAs/OLAs
- Ensure patching, hardening, backup/restore and disaster-recovery procedures are defined, scheduled, executed and evidenced — not assumed
- Manage infrastructure lifecycle: obsolescence tracking, hardware/software refresh, vendor support contracts and licences
Own the process
- Act as process owner for infrastructure Change, Incident, Problem and Configuration Management in ServiceNow; chair or co-chair the Change Advisory Board for infrastructure changes
- Enforce change control: every production change has a risk assessment, test evidence, a rollback plan and an approval trail
- Keep the CMDB accurate and complete for the infrastructure estate and make it the single source of truth
- Own major incident management for infrastructure: incident command, communication, post-incident review, and problem management through to root cause and permanent fix
Lead the team
- Lead, prioritise and develop the infrastructure team (Linux and network engineers): workload, on-call rota, skills and performance
- Set clear standards for documentation, runbooks and operational procedures, and make sure they are followed
Control and report
- Define and track operational KPIs (availability, incident volume and MTTR, change success rate, patch compliance, backup success, capacity headroom) and report regularly to managementMaintain the infrastructure risk register and drive mitigations; support security, audit and compliance reviews with evidence
Support the transition and take the target state into operations
- Be the source of truth on the current estate for the migration programme: inventory, dependencies, configurations, constraints and operational risks — and make sure migration planning is built on it
- Review migration and target designs from an operations standpoint: supportability, monitoring, backup/DR, access, network and security controls, ITSM and CMDB integration
- Define and enforce operational acceptance criteria for the AVS/EVS landscape: no workload is accepted into operations without runbooks, monitoring, backup, DR evidence, CMDB entries and trained on-call
- Prepare the team and the operating model for the target platform: skills, procedures, on-call, vendor and cloud-provider support paths
- Support cutovers from the operations side, run hyper-care, and execute on-premise decommissioning through formal change control
- Run hybrid operations during the transition without degrading service levels
Must-have experience and skills
- 8+ years in infrastructure operations, including 3+ years leading an infrastructure team in a 24/7, mission-critical or regulated environment (e.g. aviation, energy/utilities, telecom, finance, defence, healthcare)
- Strong Linux operations background (RHEL/SUSE or similar): patching, hardening, configuration management, troubleshooting
- Solid networking knowledge: TCP/IP, routing and switching, firewalls, load balancing, DNS, segmentation — enough to review designs, challenge network engineers and lead a network incident
- Virtualisation operations (VMware vSphere), storage and backup fundamentals
- Expert-level ITSM practitioner: ITIL-based operations with proven ownership of Change, Incident, Problem and Configuration Management
- Expert-level ServiceNow user and process owner: ITSM modules, CMDB, SLAs, reporting and dashboards; able to specify and validate workflow configuration with ServiceNow administrators
- Experience of moving an on-premise VMware estate to a cloud VMware platform (AVS, EVS or equivalent such as Google Cloud VMware Engine) from the operations side — discovery and dependency mapping, operational readiness, cutover support, hyper-care — and of operating the target platform afterwards
- Working knowledge of Azure or AWS core services (networking, identity, connectivity) sufficient to operate a hybrid landscape and a cloud-hosted VMware platform
- Fluent English, written and spoken; able to report clearly to senior management and non-technical stakeholders
Nice to have
- ITIL 4 Managing Professional, or ITIL 4 Foundation with practitioner modules
- ServiceNow certifications (CSA, CIS-ITSM)
- VMware VCP; Azure (AZ-104 / AZ-305) or AWS Solutions Architect; CCNA/CCNP or equivalent
- Exposure to VMware HCX and cloud-VMware networking (network extension, ExpressRoute / Direct Connect)
- Automation and observability tooling: Ansible, Terraform, Bash/Python, Zabbix/Nagios/Prometheus/Grafana, ELK
- Experience with information-security frameworks and audits (ISO 27001, NIS2) and with data-centre operations
Who you are
- Mission-critical mindset: you assume production is fragile until proven otherwise, and you plan for failure before it happens
- Meticulous and control-oriented: you check, you verify, you insist on evidence and you keep records — nothing changes in production without a plan, a review and a rollback path
- Calm and decisive under pressure; you take command of an incident rather than waiting to be asked
- Structured communicator who gives management a clear picture — including bad news, early
- Comfortable holding engineers, vendors and a migration programme to operational standards without doing their job for them
Working conditions
- Participation in the 24/7 on-call escalation rota as infrastructure incident manager
- Occasional out-of-hours work for maintenance windows and migration cutovers