Essential Duties and Responsibilities
The Hybrid Cloud Engineer supports, maintains, and improves enterprise hybrid cloud, Linux, Unix, IBM Power, AIX, storage, virtualization, automation, and resiliency platforms. This role blends hands‑on operational engineering with modern cloud practices across Microsoft Azure, Red Hat Enterprise Linux, IBM Power Systems, AIX, VIOS, HMC, NIM, SAN storage, hypervisors, Puppet, Git, Terraform, monitoring, backup/recovery, and infrastructure lifecycle management.
The engineer is expected to operate independently, execute approved changes with strong validation, troubleshoot complex incidents across infrastructure layers, produce clear documentation, and help advance secure, resilient, cost-conscious hybrid cloud operations. This role is intentionally focused on hybrid cloud, Linux/Unix, IBM Power/AIX, storage, virtualization, automation, and operational engineering.
1. IBM Power Systems and AIX Engineering – Highest Priority
- Administer IBM Power 8 and Power 9 enterprise systems, including hardware configuration, firmware lifecycle coordination, capacity planning, platform provisioning, hardware diagnostics, and system initialization support.
- Support Advanced System Management Interface (ASMI) configuration, IBM Power hardware monitoring, hardware maintenance coordination, and data center deployment activities across multiple locations.
- Administer IBM AIX operating systems in enterprise production environments, including operating system configuration, patching, upgrades, device management, logical volume management, filesystems, networking, performance tuning, security maintenance, and troubleshooting.
- Administer Virtual I/O Server (VIOS) environments, including virtual storage, virtual networking, client LPAR dependencies, resource presentation, maintenance coordination, and troubleshooting.
- Design, provision, and support LPARs, including resource allocation, storage presentation, virtual networking, capacity validation, operational standards, and lifecycle management for mission‑critical workloads.
- Administer Hardware Management Console (HMC) environments, including LPAR creation and management, system monitoring, virtualization configuration, firmware update coordination, and infrastructure issue diagnosis.
- Use Network Installation Manager (NIM) for AIX deployments, upgrades, maintenance activities, standardized builds, and disaster recovery processes.
- Use SMIT/SMITTY for AIX administration, system configuration, package/fileset maintenance, operational support, and standardized procedure development.
- Support AIX resiliency and recovery practices, including backup/restore validation, mksysb awareness, high availability support, performance review, and vendor escalation.
2. Puppet Configuration Management
- Develop, maintain, and troubleshoot Puppet‑based infrastructure automation solutions for large‑scale Unix and Linux environments.
- Create and maintain reusable Puppet modules, classes, manifests, Hiera data structures, and role/profile patterns where applicable.
- Use Puppet Enterprise or comparable Puppet control workflows to support configuration compliance, software deployment, server provisioning, patch orchestration, and baseline configuration management.
- Troubleshoot Puppet agents, catalog compilation issues, code deployments, module dependencies, environment promotion, and configuration drift.
3. Git Version Control and DevOps Practices
- Use Git and Git‑based repository platforms to manage infrastructure code, automation frameworks, configuration management content, operational scripts, and technical documentation.
- Participate in branching strategies, merge requests, conflict resolution, code reviews, release management, environment promotion, and collaborative development workflows.
- Apply DevOps best practices to infrastructure configuration, change tracking, peer review, and controlled promotion of automation changes.
4. Azure Cloud Engineering
- Design, deploy, operate, and troubleshoot Microsoft Azure infrastructure services in hybrid enterprise environments.
- Administer Azure virtual machines, networking, storage, monitoring, backup, recovery, resource governance, subscription management, tagging, policies, RBAC, and cost‑control dependencies.
- Investigate and resolve complex operational issues spanning Azure and on‑premises infrastructure while maintaining reliability, security, performance, and cost‑optimization objectives.
- Contribute to Azure architecture patterns such as landing zones, network segmentation, resiliency, resource lifecycle management, observability, and hybrid connectivity.
5. Terraform and Infrastructure as Code
- Design and maintain Infrastructure as Code solutions using Terraform for cloud and hybrid infrastructure provisioning.
- Develop reusable Terraform modules, manage state files, support environment promotion strategies, and troubleshoot Terraform deployments in enterprise cloud environments.
- Integrate Terraform workflows with Git, change management, code review, deployment pipelines, and operational validation practices.
6. Red Hat Enterprise Linux Administration
- Administer Red Hat Enterprise Linux systems in large‑scale production environments, including operating system deployment, patching, upgrades, performance tuning, package management, security hardening, lifecycle support, and troubleshooting.
- Support RHEL lifecycle planning, vulnerability remediation, repository/subscription management, SELinux awareness, Red Hat Satellite/Insights where applicable, and enterprise support processes.
- Maintain Linux configuration standards, compliance evidence, monitoring alignment, backup validation, and operational readiness.
7. Unix Shell Scripting and Automation
- Develop and maintain production‑ready automation using Bash, KornShell (ksh), and standard Unix text‑processing utilities such as awk, sed, grep, find, sort, cut, and related tools.
- Create automation for systems administration, health checks, monitoring, reporting, data processing, configuration management, maintenance activities, and operational support.
- Write maintainable scripts with logging, error handling, validation, rollback awareness, documentation, and repeatable execution patterns.
8. SAN Storage, Hypervisor, and Data Center Infrastructure
- Support enterprise SAN storage concepts and workflows, including LUN presentation, zoning coordination, host mappings, multipathing dependencies, replication awareness, capacity management, and storage‑related incident triage.
- Support virtualization and private cloud platforms such as VMware vSphere/ESXi, Nutanix AHV/HCI, IBM PowerVM, Hyper‑V, or comparable hypervisors, including VM/LPAR provisioning support, resource validation, cluster health, host maintenance, and incident troubleshooting.
- Coordinate with hardware, facilities, network, backup, storage, database, application, and vendor teams for physical and virtual server lifecycle work, data center activities, maintenance windows, and dependency validation.
9. Cross-Platform Resiliency, Troubleshooting, and Reliability Engineering
- Diagnose and resolve complex issues across physical infrastructure, operating systems, virtualization platforms, cloud services, networking, storage, automation platforms, monitoring systems, backup solutions, and application dependencies.
- Perform root cause analysis, document findings, implement corrective actions, improve operational runbooks, and reduce recurring incidents.
- Support highly available hybrid environments and help restore critical services during major incidents, maintenance events, and platform disruptions.
- Drive continuous service improvements across monitoring quality, lifecycle hygiene, vulnerability remediation, backup/recovery validation, and platform resiliency.
10. Security, Compliance, Change, and Operational Governance
- Execute scheduled vulnerability remediation, patching, operating system upgrades, firmware coordination, monitoring updates, automation releases, and approved infrastructure changes with validation and rollback awareness.
- Support PCI, SOX, audit, attack‑and‑penetration remediation, security control implementation, evidence collection, and formal change management processes.
- Produce clear SOPs, knowledge base articles, diagrams, configuration records, handoff notes, change documentation, and operational procedures.
- Collaborate effectively with cybersecurity, networking, database, application, desktop support, procurement, finance, vendors, and infrastructure teams.