Senior Cloud Infrastructure Engineer (Azure)

Numerator Career Center

India

On-site

INR 4,000,000 - 6,000,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Numerator Career Center is seeking a highly experienced Cloud Infrastructure Engineer to own the reliability of the Azure VM estate, predominantly Windows with a growing Linux footprint. You will manage patching, backups, DR, and automation through Terraform and scripting.

With 8–10 years of experience, you will build monitoring, drive auto-remediation, and lead incident analyses while documenting runbooks and standards for the platform to be operable by the whole team.

Qualifications

  • 8–10 years of cloud infrastructure experience.
  • Experience with Azure VM management (Windows/Linux).
  • Automation through Terraform and scripting.

Responsibilities

  • Own the health, reliability, and capacity of the Azure VM estate.
  • Manage patching with Azure Update Manager and maintain compliance.
  • Handle backups, DR, testing and recoverability.
  • Implement monitoring with Azure Monitor and auto-remediation.
  • Automate VM lifecycle (provision, configure, decommission) with Terraform.
  • Lead root-cause analysis for incidents affecting VM estate.
  • Collaborate with security during investigations and preserve logs.
  • Document runbooks and standards for team operability.

Skills

Terraform
Scripting
Cloud infrastructure
Incident management

Tools

Azure Monitor
KQL
Azure Update Manager
Azure Backup
Azure Site Recovery

Job description

We are seeking a highly experienced (8-10 years) Cloud Infrastructure Engineer to own the reliability of our Azure virtual machine estate. Predominantly Windows today, with a smaller Linux footprint, currently around 60 VMs and growing to 300+ as we bring more business units onto the platform. Your job is to keep it patched, backed up, monitored, and recoverable, and automate yourself out of the repetitive tasks. We expect problems to be solved once, in code, with Terraform and scripting.

Key Responsibilities
  • Own the health and reliability of the Azure VM estate (Windows and Linux), availability, performance, and capacity.
  • Own / Run patching and update management with Azure Update Manager: patch compliance, maintenance windows, remediation of failures, and handling applications that need a version held or pinned without falling out of the compliance cycle.
  • Own backup and disaster recovery: Azure Backup, Azure Site Recovery, and regular restore testing. A backup that was never restored doesn't count.
  • Build monitoring and alerting with Azure Monitor and KQL (Kusto Query Language) and drive auto-remediation so known issues fix themselves.
  • Automate VM lifecycle operations (provisioning, configuration, decommissioning) with Terraform and scripting.
  • Manage incidents affecting the VM estate, drive root cause analysis, and fix the class of problem, not just the instance.
  • Diagnose cases where the real root cause is security tooling. Antivirus/EDR flagging and quarantining an application file, for example, rather than assuming the fault sits in the application or the infrastructure.
  • Support security-led investigations on the VM estate: pull logs, process activity, and access history on request, and hold off on remediating or restarting a box until security has cleared it.
  • Reduce toil: identify repetitive manual work and eliminate it through automation.
  • Document runbooks and operational standards, so the platform is operable by the whole team.
Get your free, confidential resume review.

or drag and drop your file here.