Job Summary
We are seeking a skilled Systems Engineer Linux & Storage to join our Global Capability Center (GCC) support team. The role owns Linux infrastructure and enterprise storage environments end-to-end from proactive monitoring and automation through root cause analysis and continuous improvement – ensuring high availability, performance, and reliability of client systems. This is an operational engineering role, not a pure support function: the engineer is expected to reduce recurring issues through automation, drive standardization, and continuously improve operational practices. The candidate will act as a shift lead and primary escalation point for complex Linux and Storage incidents, driving resolution and supporting critical production environments. The primary profile we are targeting is a strong, hands-on Linux engineer with solid automation/scripting ability and practical SAN/NAS/backup experience – not necessarily SME-level depth across every storage and backup platform.
Key Responsibilities
Linux Administration
- Administer, monitor, and troubleshoot enterprise Linux environments (RHEL, CentOS, Oracle Linux, Ubuntu).
- Perform OS installation, configuration, patching, upgrades, and hardening activities.
- Manage user accounts, permissions, file systems, services, kernel parameters, and performance tuning.
- Troubleshoot system issues related to CPU, memory, storage, networking, and application performance.
- Troubleshoot cross-domain issues spanning Linux, storage, virtualization, network, and application dependencies to isolate true root cause.
- Manage packages and repositories using patch/repository management tooling (e.g., Satellite, Spacewalk, YUM/DNF/APT repos).
- Use Git/version control for scripts, automation content, and configuration management.
Automation & Scripting
- Drive automation as a core responsibility, not an optional add-on – with strong hands-on Shell/Bash scripting and Ansible for configuration management, patching, and repetitive operational tasks.
- Build and maintain Ansible playbooks/roles for provisioning, patching, hardening, and routine operational workflows.
- Python scripting is a desirable addition for extending automation and tooling capabilities.
- Apply an automation/AI mindset – proactively using modern AI-assisted tools for scripting, troubleshooting, log analysis, and documentation to improve speed and quality of resolution.
- Identify recurring issues and build automated, permanent fixes to reduce repeat incidents and manual effort.
Storage Administration
- Manage and support enterprise SAN/NAS storage platforms with meaningful, hands-on experience – deep expertise across every storage vendor/product is not required.
- Provision storage, create and manage LUNs, storage pools, RAID groups, volumes, and snapshots.
- Monitor storage capacity, utilization, performance, and availability.
- Troubleshoot storage connectivity issues involving Fibre Channel (FC), iSCSI, multipathing, and zoning.
- Support storage migrations and expansion activities.
Backup & Recovery
- Administer enterprise backup solutions such as Veritas NetBackup, Commvault, Veeam, or similar, with practical hands-on experience on at least one major platform.
- Perform backup monitoring, restoration requests, disaster recovery validations, and troubleshooting.
- Ensure compliance with backup policies and recovery objectives.
Security, Patching & Compliance
- Own the end-to-end patching lifecycle – planning, testing, scheduling, and execution – across Linux estates.
- Drive vulnerability remediation in coordination with security teams, tracking findings through to closure within SLA.
- Implement and maintain security hardening baselines and compliance standards (e.g., CIS benchmarks) across Linux systems.
- Support audit and compliance reporting for patch and vulnerability status.
Operational Engineering & Continuous Improvement
- Proactively monitor systems and storage to identify risks and degradation before they cause incidents.
- Lead root cause analysis (RCA) for critical incidents and drive permanent, automated fixes rather than workarounds.
- Track and reduce recurring issues through problem management and process improvement.
- Continuously improve monitoring, automation, and operational runbooks to raise efficiency and reliability.
Incident & Operations Support
- Act as Shift Lead for Linux/Storage operations and support teams.
- Lead complex incident triage, root cause analysis, and resolution activities.
- Coordinate with cross-functional teams, vendors, and customers during major incidents.
- Participate in on-call support and rotational shift schedules.
- Prepare incident reports, technical documentation, SOPs, and knowledge articles.
Required Skills
Linux
- Strong expertise in:
- Red Hat Enterprise Linux (RHEL)
- CentOS / Oracle Linux / Ubuntu
- System performance tuning
- Filesystem management (XFS, EXT4, LVM)
- Patch and repository management tooling
Automation & Scripting
- Hands-on experience with:
- Shell/Bash scripting (required, hands-on)
- Ansible for configuration management and automation (required, hands-on)
- Git/version control for scripts and automation content
- Python scripting (desirable)
Storage
- Hands‑on, practical experience with:
- SAN & NAS technologies
- Fibre Channel and iSCSI environments
- Storage provisioning and capacity management
- Exposure to one or more platforms such as EMC Dell, NetApp, Hitachi, IBM, or Pure Storage (broad multi‑vendor SME depth not required)
Backup
- Practical experience with at least one enterprise backup tool:
- Veritas NetBackup
- CommvaultVeeam
- Backup Exec (preferred)
Security & Compliance
- Working knowledge of patching, vulnerability remediation workflows, and security hardening practices.
- Familiarity with compliance frameworks/benchmarks (e.g., CIS) for Linux systems.
Additional Skills
- Understanding of ITIL processes (Incident, Problem, Change Management).
- Experience supporting large-scale enterprise production environments with shift/on-call responsibilities.
- Strong analytical, troubleshooting, and communication skills.
- Experience leading technical bridge calls and handling critical incidents.
- Comfortable using AI-assisted tools to speed up scripting, troubleshooting, log analysis, and documentation.
Preferred Qualifications
- Relevant certifications are highly desirable:
- RHCSA / RHCE
- Red Hat Specialist Certifications
- NetApp, Dell EMC, or Storage Vendor Certifications
- ITIL Foundation
- Working knowledge/exposure to Docker and Kubernetes.
- Exposure to Linux integration with Active Directory (SSSD, Kerberos, keytabs).
Success Profile
- Acts as the primary technical resource for Linux and Storage escalations.
- Independently handles critical production incidents.
- Drives automation and process improvement to reduce manual effort and recurring issues.
- Provides mentoring and guidance to junior engineers.
- Demonstrates strong ownership, customer focus, and operational excellence.