CLOUD ARCHITECT - Ansible

Happiest Minds Technologies

Bengaluru

On-site

INR 1,400,000 - 2,100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Happiest Minds Technologies seeks a Cloud Disaster Recovery Engineer to design, implement, and maintain DR solutions across cloud and on‑prem environments. You will work on high availability architectures, automation, and failover processes using Terraform, Ansible, PowerShell, and Python.

You will collaborate with IT ops, cybersecurity, and business continuity teams to enhance resilience, perform failure mode analysis, and lead DR drills and validation efforts.

Qualifications

  • Bachelor's degree in computer science, information technology, cybersecurity, or related field.
  • 5+ years in IT infrastructure, cloud engineering, disaster recovery, or resilience engineering.
  • Expertise in disaster recovery planning and HA design.
  • Hands-on with cloud resilience strategies (AWS/Azure/GCP) and cloud-native DR tools.
  • Experience automating AWS DR using Boto3 and cross-region failover.

Responsibilities

  • Design, implement, and maintain HA and fault-tolerant architectures across cloud and on-prem environments.
  • Develop and maintain DR solutions ensuring defined RTO/RPO.
  • Automate DR and failover using IaC and scripting (Terraform, Ansible, PowerShell, Python).
  • Integrate resilience into system design, deployment, and operations.
  • Perform failure mode analysis to address vulnerabilities.
  • Design automated DR runbooks and post-failover health checks.
  • Develop and execute DR drills, validate backups and replication.
  • Monitor failover performance and optimize configurations.
  • Lead incident response during disruptions and coordinate with cyber resilience teams.
  • Produce reports on resilience testing results and risk mitigation.

Skills

AWS
Azure
GCP
Terraform
Ansible
PowerShell
Python
Boto3
EC2
RDS
S3
Route53
IAM
AWS Backup
EDRS

Education

Bachelor's degree in Computer Science or related field

Tools

Veeam
Commvault
Zerto
Azure Site Recovery

Job description

The Cloud Disaster Recovery Engineer is responsible for designing, implementing, and maintaining Disaster Recovery (DR) solutions to ensure the organization's technology infrastructure and critical systems can withstand and recover from disruptions. This role involves hands-on work with high availability (HA) architectures, disaster recovery strategies, automation, and failover solutions in on-premises, cloud, and hybrid environments.

The ideal candidate will have expertise in IT resilience, infrastructure engineering, and cloud-based recovery solutions, working closely with IT operations, cybersecurity, and business continuity teams to enhance the organization's overall technology resilience posture.

Technology Resilience & Disaster Recovery Engineering

Design, implement, and maintain highly available (HA) and fault-tolerant architecture across cloud (AWS, Azure, GCP) and on-premises environments.

Develop and maintain disaster recovery (DR) solutions, ensuring that IT systems meet defined recovery time objectives (RTO) and recovery point objectives (RPO).

Implement automation and orchestration for disaster recovery and failover processes using Infrastructure as Code (IaC) and scripting tools (Terraform, Ansible, PowerShell, Python).

Work with IT infrastructure and application teams to integrate resilience best practices into system design, deployment, and operations.

Perform failure mode analysis (FMA) to identify and address system vulnerabilities.

Disaster Recovery Testing & Validation

Design and implement fully automated disaster recovery runbooks using Ansible and Python using one-click or event-triggered failover systems.

Develop automated recovery verification (post-failover health checks).

Develop and execute disaster recovery drills and failover testing, identifying gaps and improvements.

Automate DR testing, validation, and reporting.

Conduct regular validation of backup and replication strategies, ensuring data integrity and availability.

Monitor system failover and recovery performance, optimizing configurations to improve response times.

Incident Response & Crisis Management Support

Act as a technical lead during disruptions and disaster recovery events, ensuring rapid system recovery.

Work closely with cybersecurity teams to integrate DR solutions with cyber resilience strategies, ensuring quick restoration from ransomware or cyberattacks.

Support post-incident analysis and recommend improvements to resilience strategies.

Monitoring, Compliance & Reporting

Implement and maintain resilience monitoring tools, ensuring continuous tracking of system availability and DR readiness.

Ensure compliance with industry standards and regulatory requirements(e.g., ISO 27001, NIST, FFIEC, SOC2).

Provide technical input for audits and regulatory assessments related to technology resilience.

Generate reports on resilience testing results, failover performance, and risk mitigation efforts.

Collaboration & Training

Work closely with IT teams, business continuity professionals, and cloud architects to ensure resilience strategies align with business needs.

Provide training and technical guidance to IT staff on disaster recovery best practices and system failover configurations.

Assist in the development of technical documentation and playbooks for disaster recovery and resilience processes.

Qualifications & Experience

Bachelor's degree in computer science, Information Technology, Cybersecurity, or a related field.

5+ years of experience in IT infrastructure, cloud engineering, disaster recovery, or resilience engineering.

Expertise in disaster recovery planning and high-availability (HA) solution design.

Hands-on experience with cloud resilience strategies (AWS, Azure, or GCP) and cloudnative DR tools.

Hands-on experience automating AWS disaster recovery using:

Boto3 for orchestration of EC2, RDS, S3, Route 53, IAM, AWS Backup, and EDRS.

Cross-region replication and failover strategies

AWS-native DR patterns (pilot light, warm standby, multi-region active/active)

Automation of AMI lifecycle, backup validation, and restore testing

Advanced automation engineering experience, including:

Designing and maintaining enterprise-scale Ansible automation frameworks (roles, collections, dynamic inventories, Ansible Automation Platform/AWX)

Developing production-grade Python automation using Boto3

Buildingevent-driven automation workflows for failover and recovery

Implementing idempotent Infrastructure as Code (IaC) patterns

Integrating automation into CI/CD pipelines (e.g., GitHub Actions, Jenkins)

Experience with backup, replication, and data protection solutions (e.g., Veeam, Commvault, Zerto, Azure Site Recovery)

Knowledge of networking, storage, virtualization, and hybrid-cloud architecture.

Cloud Automation and orchestration at scale is a high priority for this position. Candidates must be bringing extensive demonstrable experience in coding and scripting of cloud resources in AWS and other cloud environments.

Certifications

AWS Certified Solutions Architect,

Red Hat Certified Engineer,

Microsoft Azure Administrator,

Certified Business Continuity Professional (CBCP), or

Disaster Recovery Certified Specialist (DRCS)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead, Resiliency Engineer | Pune
Lead, Resiliency Engineer | Pune

Northern Trust • Pune District, Bengaluru

Hybrid
INR 3,000,000 - 4,500,000
AWS Cloud Engineer
AWS Cloud Engineer

VMC Soft Technologies, Inc • Chennai District

On-site
INR 1,200,000 - 1,600,000
IT Resiliency Engineer
IT Resiliency Engineer

AppSierra • Pune City

On-site
INR 1,500,000 - 2,500,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Healthedge • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Ansible Automation Engineer
Ansible Automation Engineer

PwC India • Pune District

On-site
INR 1,200,000 - 2,500,000
Infrastructure and Platform Engineer Architect
Infrastructure and Platform Engineer Architect

PwC India • Pune District

On-site
INR 2,000,000 - 4,000,000
Associate Architect - Site Reliability
Associate Architect - Site Reliability

Highradius • Hyderabad

On-site
INR 2,500,000 - 4,200,000
Cloud Automation Engineer
Cloud Automation Engineer

Sunovaa Tech • Bengaluru

On-site
INR 1,500,000 - 2,100,000
AWS Cloud Specialist
AWS Cloud Specialist

Coforge • Greater Noida

On-site
INR 900,000 - 1,300,000
Cloud Resiliency Engineer
Cloud Resiliency Engineer

Concentrix • India

On-site
INR 1,200,000 - 1,800,000