The L3 Splunk Admin & IT Operations Lead - RTM is responsible for providing high-level operational and technical leadership to Airbus. In addition to expert-level Splunk administration, this role acts as the operational backbone for the existing support team - driving day-to-day IT Operations Management (ITOM), continual system reliability, and overall infrastructure availability. Working in a self-sufficient, multi-disciplinary team, the individual will mentor L1/L2 engineers, manage critical operational issues, and ensure strict alignment with ITIL processes (Incident, Problem, and Change Management).
Key Deliverables
- Deliver business value and maintain high operational reliability.
- Provide strong operational and technical leadership to elevate the performance and capability of the L1/L2 support team.
- Provide transparent updates during Agile rituals (remaining work, roadblocks) and elevate risks proactively.
- Consistently ensure on-time, high-quality (first-time-right) delivery of operational sprints and platform enhancements.
- Introduce innovative, cost-effective automation and operational optimization solutions.
- Aim for high customer satisfaction by maintaining SLA/OLA compliance and minimizing mean time to resolution (MTTR).
Key Responsibilities
- Take full ownership of technical operations and be available 24x7 for crisis management and major incidents.
- Provide daily operational guidance, mentoring, and escalation support to the L1/L2 operations teams.
- Perform trend analysis on daily, weekly, and monthly KPI reports to identify potential operational risks and execute continuous improvement plans.
- Review business-specific KPI dashboards and implement value-added monitoring and operational metrics.
- Drive Incident, Problem, and Change Management workflows to ensure resolution strictly within defined SLAs and OLAs.
- Conduct Root Cause Analysis (RCA) for complex outages and lead preventative action plans for permanent issue resolution.
- Maintain operational excellence through technical audits, system optimization, and standardizing team workflows.
- Create, update, and review Knowledge Base articles, standard operating procedures (SOPs), and deliver training sessions to support staff.
- Collaborate with the Product Owner and IT leadership to align operational capabilities with future technical innovations.
- Drive end-to-end automation of recurring tasks using DevOps tools to increase efficiency across RTM run operations.
Knowledge & Skills
Engineering graduate with 8-10 years of experience in Infrastructure Monitoring, Splunk Administration, and IT Operations Management in a matrix organization.
Must Know (Expert Level)
- Strong technical ownership with a proven track record of guiding L1/L2 support teams and resolving complex operational issues.
- Splunk Infrastructure management, multi-site clustering, log deployment, integration, and administration (Splunk Certification preferred).
- Strong IT Operations Management (ITOM) expertise, including SLA/OLA management, shift alignment, and escalation protocols.
- ITIL framework execution (Incident, Change, Problem, and Escalation Management).
- Certified Red Hat Linux Administrator with strong practical knowledge of Windows OS environments.
- Cloud Infrastructure: AWS (EC2, S3, Route53, IAM).
- Automation & Scripting: Shell, Python.
- Networking Fundamentals: DHCP, DNS, SSL/TLS certifications, and traffic routing.
- DevOps Tools: Git, Ansible, Jenkins.
- Identity & Access Management: SAML, OAuth integrations.
- System performance troubleshooting and RCA methodologies.
- Excellent verbal and written communication skills.
Intermediate Knowledge
- Agile / Scrum ways of working.
- ServiceNow (ITSM platform administration and workflow usage).
Basic Knowledge
- JIRA, Confluence, VersionOne, or equivalent project tracking tooling.
This job requires an awareness of any potential compliance risks and a commitment to act with integrity, as the foundation for the Company's success, reputation and sustainable growth.