Cloud Disaster Recovery Engineer

Steady Rabbit

United States

Hybrid

USD 120,000 - 180,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Steady Rabbit seeks a Cloud Disaster Recovery Engineer to design, implement, and operate DR solutions across cloud, on‑prem, and hybrid environments. You will build highly available architectures and automate failover processes using IaC and scripting languages.

The role requires 5+ years in IT infrastructure, cloud engineering, and disaster recovery, with hands‑on experience in AWS, Azure, or GCP and strong scripting skills.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology or related field.
  • 5+ years of IT infrastructure, cloud engineering, disaster recovery, or resilience engineering.
  • Expertise in disaster recovery planning and high-availability solution design.
  • Hands-on experience with cloud resilience strategies (AWS/Azure/GCP) and cloud-native DR tools.
  • Strong automation engineering experience with Ansible, Python, Terraform, and CI/CD.

Responsibilities

  • Design, implement, and maintain highly available and fault-tolerant architectures across cloud and on-prem environments.
  • Develop and maintain disaster recovery solutions with defined RTOs and RPOs.
  • Automate DR and failover processes using IaC and scripting (Terraform, Ansible, PowerShell, Python).
  • Collaborate with IT operations, cybersecurity, and business continuity teams on resilience.
  • Perform failure mode analysis to identify vulnerabilities.

Skills

IT infrastructure
Cloud engineering
Disaster recovery planning
Automation engineering
Infrastructure as Code
SRE mindset

Education

Bachelor's degree in CS/IT/Cybersecurity

Tools

AWS
Azure
GCP
Terraform
Ansible
PowerShell
Python
Boto3
GitHub Actions
Jenkins
AWX
Veeam
Azure Site Recovery

Job description

Cloud Disaster Recovery Engineer

Experience: 5+ years

Role Overview

The Cloud Disaster Recovery Engineer is responsible for designing, implementing, and maintaining Disaster Recovery (DR) solutions to ensure the organization's technology infrastructure and critical systems can withstand and recover from disruptions. This role involves hands-on work with high availability (HA) architectures, disaster recovery strategies, automation, and failover solutions in on-premises, cloud, and hybrid environments.

The ideal candidate will have expertise in IT resilience, infrastructure engineering, and cloud-based recovery solutions, working closely with IT operations, cybersecurity, and business continuity teams to enhance the organization's overall technology resilience posture.

Key Responsibilities
Technology Resilience & Disaster Recovery Engineering
  • Design, implement, and maintain highly available (HA) and fault-tolerant architectures across cloud (AWS, Azure, GCP) and on-premises environments.
  • Develop and maintain disaster recovery (DR) solutions, ensuring that IT systems meet defined recovery time objectives (RTO) and recovery point objectives (RPO).
  • Implement automation and orchestration for disaster recovery and failover processes using Infrastructure as Code (IaC) and scripting tools (Terraform, Ansible, PowerShell, Python).
  • Work with IT infrastructure and application teams to integrate resilience best practices into system design, deployment, and operations.
  • Perform failure mode analysis (FMA) to identify and address system vulnerabilities.
Disaster Recovery Testing & Validation
  • Design and implement fully automated disaster recovery runbooks using Ansible and Python using one-click or event-triggered failover systems.
  • Develop automated recovery verification (post-failover health checks).
  • Develop and execute disaster recovery drills and failover testing, identifying gaps and improvements.
  • Automate DR testing, validation, and reporting.
  • Conduct regular validation of backup and replication strategies, ensuring data integrity and availability.
  • Monitor system failover and recovery performance, optimizing configurations to improve response times.
Incident Response & Crisis Management Support
  • Act as a technical lead during disruptions and disaster recovery events, ensuring rapid system recovery.
  • Work closely with cybersecurity teams to integrate DR solutions with cyber resilience strategies, ensuring quick restoration from ransomware or cyberattacks.
  • Support post-incident analysis and recommend improvements to resilience strategies.
Monitoring, Compliance & Reporting
  • Implement and maintain resilience monitoring tools, ensuring continuous tracking of system availability and DR readiness.
  • Ensure compliance with industry standards and regulatory requirements (e.g., ISO 27001, NIST, FFIEC, SOC 2).
  • Provide technical input for audits and regulatory assessments related to technology resilience.
  • Generate reports on resilience testing results, failover performance, and risk mitigation efforts.
Collaboration & Training
  • Work closely with IT teams, business continuity professionals, and cloud architects to ensure resilience strategies align with business needs.
  • Provide training and technical guidance to IT staff on disaster recovery best practices and system failover configurations.
  • Assist in the development of technical documentation and playbooks for disaster recovery and resilience processes.
Qualifications & Experience
  • Bachelor's degree in Computer Science, Information Technology, Cybersecurity, or a related field.
  • 5+ years of experience in IT infrastructure, cloud engineering, disaster recovery, or resilience engineering.
  • Expertise in disaster recovery planning and high-availability (HA) solution design.
  • Hands-on experience with cloud resilience strategies (AWS, Azure, or GCP) and cloud-native DR tools.
  • Hands-on experience automating AWS disaster recovery using:
    • Boto3 for orchestration of EC2, RDS, S3, Route 53, IAM, AWS Backup, and EDRS.
    • Cross-region replication and failover strategies.
    • AWS-native DR patterns (pilot light, warm standby, multi-region active/active).
    • Automation of AMI lifecycle, backup validation, and restore testing.
  • Advanced automation engineering experience, including:
    • Designing and maintaining enterprise-scale Ansible automation frameworks (roles, collections, dynamic inventories, Ansible Automation Platform/AWX).
    • Developing production-grade Python automation using Boto3.
    • Building event-driven automation workflows for failover and recovery.
    • Implementing idempotent Infrastructure as Code (IaC) patterns.
    • Integrating automation into CI/CD pipelines (e.g., GitHub Actions, Jenkins).
  • Experience with backup, replication, and data protection solutions (e.g., Veeam, Commvault, Zerto, Azure Site Recovery).
  • Knowledge of networking, storage, virtualization, and hybrid-cloud architectures.
  • Cloud automation and orchestration at scale is a high priority for this position. Candidates must bring extensive demonstrable experience in coding and scripting of cloud resources in AWS and other cloud environments.
Certifications
  • AWS Certified Solutions Architect
  • Red Hat Certified Engineer
  • Microsoft Azure Administrator
  • Certified Business Continuity Professional (CBCP), or
  • Disaster Recovery Certified Specialist (DRCS)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Operations Engineer (Cloud)
Operations Engineer (Cloud)

Stifel Financial Corp. • Memphis (TN)

On-site
USD 90,000 - 120,000
Senior Disaster Recovery & Resilience Architect
Senior Disaster Recovery & Resilience Architect

Unisys • Norwich (CT)

On-site
USD 140,000 - 190,000
Cloud Disaster Recovery Architect
Cloud Disaster Recovery Architect

Steady Rabbit • United States

Hybrid
USD 120,000 - 180,000
Disaster Recovery
Disaster Recovery

ManpowerGroup Global, Inc. • Charlotte (NC), Town of Norway (WI)

Hybrid
USD 82,656 - 96,432
Medical and Prescription Drug Plans
Dental Plan
Vision Plan
+8
Senior Technology Resiliency and Recovery Engineer
Senior Technology Resiliency and Recovery Engineer

Jobtailor • Vienna (VA)

On-site
USD 140,000 - 210,000
Technical Analyst
Technical Analyst

The Judge Group • Tempe (AZ)

On-site
USD 120,000 - 160,000
Disaster Recovery Specalist
Disaster Recovery Specalist

Insight Global • New York (NY)

On-site
USD 90,000 - 120,000
Senior Principal Disaster Recovery Architecture – Engineering
Senior Principal Disaster Recovery Architecture – Engineering

Jobtailor • Village of Sleepy Hollow (NY)

On-site
USD 180,000 - 240,000
Cloud Disaster Recovery Program Manager 3638404
Cloud Disaster Recovery Program Manager 3638404

Axiom Path • Town of Charlotte (NY)

Hybrid
Cloud Engineer
Cloud Engineer

Veriipro • Plano (TX)

On-site
USD 120,000 - 140,000