We are looking for an experienced Cloud Infrastructure Manager to oversee and manage our enterprise IT infrastructure across Azure, Google Cloud Platform (GCP), and on-premises environments. The ideal candidate will ensure high availability, security, scalability, cost optimization, and operational excellence while leading infrastructure initiatives and supporting business‑critical applications.
Key Responsibilities:
- Administer, maintain, and optimize Ubuntu/Linux servers across Azure, GCP, and on-premises environments.
- Manage cloud infrastructure including Azure Virtual Machines, Google Compute Engine, networking, storage, IAM, RBAC, VPNs, firewalls, and hybrid connectivity.
- Automate infrastructure provisioning and configuration using Terraform, Ansible, Bash, Python, Azure CLI, and Google Cloud CLI.
- Implement infrastructure security through patch management, system hardening, vulnerability remediation, access controls, and secrets management.
- Maintain monitoring, logging, alerting, backup, disaster recovery, and business continuity solutions using industry-standard tools.
- Support CI/CD pipelines, Docker-based deployments, and collaborate with DevOps, AI/ML, GIS, Data Engineering, and Application teams.
- Perform capacity planning, cloud cost optimization, resource rightsizing, and infrastructure performance tuning.
- Lead incident management, root cause analysis (RCA), production support, and infrastructure change management.
- Maintain infrastructure documentation, SOPs, architecture diagrams, inventories, and operational standards.
- Lead infrastructure engineers, coordinate with vendors, cloud providers, and ensure compliance with security and operational policies.
Preferred Skills :
- 4+ years of Strong hands‑on experience with Ubuntu/Linux administration in production environments.
- Expertise in Microsoft Azure and Google Cloud Platform (GCP).
- Strong knowledge of networking (TCP/IP, DNS, VPN, NAT, Load Balancers, VPC, VNet, Firewalls).
- Hands‑on experience with Terraform, Ansible, Infrastructure as Code (IaC), and automation.
- Experience with Linux security, IAM, RBAC, SSH, secrets management, and system hardening.
- Experience with monitoring and observability tools such as Azure Monitor, Google Cloud Monitoring, Prometheus, Grafana, ELK/OpenSearch, and Alert manager.
- Working knowledge of Docker and CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI.
- Experience with backup, disaster recovery, business continuity planning, and cloud cost optimization.
- Exposure to PostgreSQL, MySQL, Redis, Kafka, Cassandra, or similar infrastructure technologies is an added advantage.
- Excellent troubleshooting, documentation, leadership, vendor management, and stakeholder communication skills.