Cloud Infrastructure Lead

CPR Vision Management

Malaysia

On-site

MYR 180,000 - 300,000

Full time

21 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

CPR Vision is seeking a Cloud Infrastructure Lead to own Microsoft Azure environments and production infrastructure. This hands-on role focuses on operating, improving, and securing our cloud platform, coordinating managed-service providers and internal teams to address risks and operational actions through to completion.

You will manage day-to-day Azure operations, implement IaC pipelines, monitor performance, and contribute to disaster recovery and cost management.

Qualifications

  • 7+ years of experience in infrastructure, cloud operations or systems engineering.
  • 3+ years of hands-on experience with Azure production environments.
  • Strong knowledge of Azure networking, identity, monitoring, backup and governance.
  • Experience in incident management and infrastructure troubleshooting.
  • Experience with Windows and/or Linux administration.
  • Experience working with managed-service providers or vendors.
  • Ability to coordinate across teams and document actions to resolution.

Responsibilities

  • Own day-to-day administration and health of CPR Vision’s Azure environments.
  • Maintain subscriptions, resource groups, access controls, policies, networking and backups.
  • Support Azure services like App Service, Azure SQL, Storage, VMs, Key Vault, API Management, Monitor, Log Analytics and Application Insights.
  • Manage Azure networking components (VNets, subnets, NSGs, VPNs, Private Endpoints, App Gateway, WAF).
  • Coordinate incident escalation, root-cause analysis and follow-up actions with teams and providers.
  • Promote secure CI/CD pipelines and IaC practices; integrate security checks in delivery pipelines.
  • Monitor production infrastructure availability, performance and capacity; document runbooks and procedures.
  • Coordinate disaster-recovery, backup testing and data-residency considerations.

Skills

Azure
Cloud security
Incident management
DevOps practices
Vendor coordination
Documentation
Stakeholder management

Education

Bachelor's degree in Computer Science / IT

Tools

Terraform
Pulumi
PowerShell
Bash
Azure CLI
Kubernetes basics

Job description

The Cloud Infrastructure Lead will be CPR Vision’s primary internal owner for Microsoft Azure and production infrastructure.

This is a hands-on technical role rather than a purely managerial position. You will operate and improve our Azure environments, troubleshoot infrastructure issues, coordinate managed-service providers and ensure that infrastructure risks and operational actions are properly addressed.

You are not expected to be a specialist in every security, audit and regulatory domain. CPR Vision works with Volaris Group teams, managed-service providers, auditors and specialist vendors. However, you must understand the infrastructure implications, coordinate the relevant parties and take ownership of actions through to completion.

Azure and Cloud Infrastructure
  • Own the day-to-day administration and operational health of CPR Vision’s Microsoft Azure environments.
  • Maintain Azure subscriptions, resource groups, access controls, policies, networking, monitoring, backups and production configurations.
  • Support Azure services including App Service, Azure SQL, Storage, Virtual Machines, Key Vault, API Management, Azure Monitor, Log Analytics and Application Insights.
  • Manage Azure networking components such as VNets, subnets, NSGs, VPNs, Private Endpoints, Application Gateway and WAF.
  • Manage infrastructure identity and access through Microsoft Entra ID and role-based access control.
  • Review the existing Azure environment and recommend practical improvements to security, reliability, manageability and cost.
  • Support new customer environments, migrations, production releases and market launches.
  • Maintain accurate cloud architecture diagrams, asset records, access documentation and operating procedures.
DevOps, Automation and Release Enablement
  • Partner with software-development teams to design, operate and improve secure CI/CD pipelines for applications, databases and infrastructure.
  • Promote disciplined source-control and release practices, including pull requests, deployment approvals, separation of duties, environment controls and traceability of production changes.
  • Support appropriate deployment strategies, including rolling, blue/green or canary approaches, where they reduce production risk.
  • Integrate infrastructure and security checks into delivery pipelines where practical, such as IaC scanning, secret detection, dependency or container-image scanning and policy validation.
  • Improve observability across applications and infrastructure, and use deployment frequency, change-failure rate, lead time and recovery time to guide continuous improvement.
  • Automate repetitive operational work with PowerShell, Bash, Azure CLI, Pulumi or other suitable tooling while maintaining clear ownership, documentation and supportability.
Infrastructure Operations and Reliability
  • Monitor the availability, performance and capacity of production infrastructure.
  • Investigate and resolve infrastructure incidents, working with developers and managed-service providers where necessary.
  • Coordinate incident escalation, root-cause analysis and follow-up actions.
  • Maintain infrastructure runbooks, escalation procedures and recovery documentation.
  • Coordinate patching, maintenance, certificate renewals, backups and infrastructure upgrades.
  • Identify recurring problems and implement sustainable fixes.
  • Support an appropriate escalation arrangement for critical production incidents.
Backup, Disaster Recovery and Business Continuity
  • Maintain backup and recovery configurations for priority systems.
  • Work with stakeholders to document recovery time and recovery point requirements.
  • Coordinate and document disaster-recovery and backup-restoration testing.
  • Identify infrastructure single points of failure and recommend practical improvements.
  • Keep recovery procedures accurate, accessible and regularly tested.
Managed-Service and Vendor Coordination
  • Act as CPR Vision’s primary technical contact for cloud, hosting, network and infrastructure service providers.
  • Monitor vendor performance against agreed service levels and escalation timelines.
  • Review vendor recommendations, incident findings and remediation plans.
  • Coordinate regular operational reviews covering incidents, vulnerabilities, capacity, costs and outstanding actions.
  • Escalate material risks, delays or recurring service issues to CPR Vision leadership.
Infrastructure Security
  • Ensure required endpoint, server and cloud-security tools are correctly deployed and operational.
  • Coordinate infrastructure-security activities involving CrowdStrike, Rapid7, Microsoft Defender for Cloud and other approved technologies.
  • Review vulnerability findings and coordinate remediation with developers, vendors and Volaris Security.
  • Maintain clear vulnerability-remediation priorities and track actions through to closure.
  • Support investigations into infrastructure and endpoint-security alerts.
  • Strengthen access controls, patching, secrets management, logging, backups and environment separation.
  • Support customer security assessments, audit requests and evidence collection relating to infrastructure.
  • The role will coordinate with specialist security teams and vendors and will not be solely responsible for operating a full security operations centre.
Alibaba Cloud and China-Hosted Environments
  • Oversee the operational health of CPR Vision’s Alibaba Cloud and China-hosted environments.
  • Coordinate local hosting and managed-service partners.
  • Maintain documentation covering infrastructure, access, monitoring, backups and system dependencies.
  • Work with internal and external specialists to ensure hosting and data-residency requirements are followed.
  • Include China-hosted environments within CPR Vision’s incident, vulnerability, backup and disaster-recovery processes.
  • Deep China regulatory expertise is advantageous but is not a mandatory requirement.
Cost and Capacity Management
  • Monitor cloud expenditure and investigate unexpected cost changes.
  • Maintain Azure budgets, alerts, tagging and basic cost-allocation practices.
  • Identify unused resources and practical opportunities for right-sizing or storage optimisation.
  • Produce concise monthly reporting covering infrastructure health, costs, incidents, vulnerabilities and risks.
  • Work with finance, leadership and vendors on infrastructure forecasts and renewals.
Internal IT Oversight
  • Coordinate employee onboarding and offboarding from an IT-access and equipment perspective.
  • Maintain accurate hardware, software and technology-asset records.
  • Coordinate account administration and employee IT requirements with Volaris IT and external providers.
  • Ensure endpoint-security, encryption and patching requirements are applied to company devices.
  • Provide escalation support for complex internal IT issues.
  • Routine first-line employee support may be delivered by an internal administrator or external support provider. This role is responsible for oversight and escalation rather than acting primarily as the company helpdesk.
What Success Looks Like

Within the first 6–12 months, the successful candidate will have:

  • Established clear visibility and ownership of CPR Vision’s infrastructure estate.
  • Improved Azure documentation, access controls, monitoring and operational standards.
  • Strengthened incident escalation and reduced repeat infrastructure issues.
  • Completed recovery testing for priority customer platforms.
  • Improved managed-service-provider accountability.
  • Identified practical opportunities to optimise cloud expenditure.
  • Documented key operational risks and implemented a prioritised improvement plan.
  • Improved coordination between infrastructure, development, security and client-delivery teams.
Required Experience
  • Approximately 7 or more years of experience in infrastructure, cloud operations, DevOps or systems engineering.
  • At least 3 years of meaningful hands-on experience supporting Microsoft Azure production environments.
  • Strong working knowledge of Azure networking, identity, monitoring, backup and governance.
  • Experience managing production incidents and troubleshooting infrastructure issues.
  • Working knowledge of Windows and/or Linux administration.
  • Familiarity with SQL Server, Azure SQL or MySQL infrastructure operations.
  • Experience working with managed-service providers or technology vendors.
  • Familiarity with endpoint security, vulnerability management and infrastructure-security practices.
  • Ability to understand technical risks and coordinate actions across multiple teams.
  • Good documentation, communication and stakeholder-management skills.
  • Strong personal ownership and willingness to follow issues through to resolution.
Advantageous Experience
  • Terraform, Pulumi, Bicep, PowerShell or Bash.
  • Alibaba Cloud or other public-cloud platforms.
  • CrowdStrike, Rapid7 or Microsoft Defender.
  • API Management, Application Gateway, WAF or Key Vault.
  • Containerised workloads or basic Kubernetes exposure.
  • SOC 2, ISO 27001 or customer-security assessments.
  • Cloud-cost optimisation and capacity planning.
  • China-hosted systems or data-residency requirements.
Preferred Certifications
  • One or more of the following would be advantageous:
  • Microsoft Certified: Azure Solutions Architect Expert
  • Microsoft Certified: Azure Security Engineer Associate
  • HashiCorp Certified: Terraform Associate
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Confidential Jobs • Kuala Lumpur

On-site
MYR 100,000 - 160,000
Cloud Infrastructure Engineering, Principal
Cloud Infrastructure Engineering, Principal

AIA Malaysia • Kuala Lumpur

On-site
MYR 180,000 - 320,000
Senior Cloud Engineer
Senior Cloud Engineer

All jobs • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Cloud Infrastructure Engineering Principal
Cloud Infrastructure Engineering Principal

AIA Malaysia • Kuala Lumpur

On-site
MYR 240,000 - 320,000
Azure Cloud Infra Lead: DevOps, Security & Reliability
Azure Cloud Infra Lead: DevOps, Security & Reliability

CPR Vision Management • Malaysia

On-site
MYR 180,000 - 300,000
Engineer, IT (Cloud)
Engineer, IT (Cloud)

FFM Berhad • Sungai Buloh

On-site
MYR 60,000 - 90,000
Cloud Engineer (AWS & Multi-Cloud) - Cigna Healthcare
Cloud Engineer (AWS & Multi-Cloud) - Cigna Healthcare

cigna • Kuala Lumpur

Hybrid
MYR 120,000 - 240,000
Senior Cloud Engineer
Senior Cloud Engineer

Logicalis Asia Pacific • Kuala Lumpur

On-site
MYR 90,000 - 120,000
Senior Executive Infrastructure and Cloud Engineering
Senior Executive Infrastructure and Cloud Engineering

Tune Protect Group • Kampung Malaysia Tambahan

Hybrid
MYR 120,000 - 190,000
Senior Infrastructure & Cloud
Senior Infrastructure & Cloud

Mode Fair Sdn Bhd • Kuala Lumpur

On-site
MYR 240,000 - 360,000
Competitive benefits
On-site at Sunway Tower