Cloud Resiliency Engineer - Automation & DR Leader

KeyBank

Brooklyn (OH)

On-site

USD 63,000 - 96,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

KeyBank is seeking a Resiliency Engineer to design, build, and improve reliability, availability, and recoverability across on-premises, hybrid, and cloud environments. This hands-on role requires coding skills and a capacity to operate during live recovery exercises, coordinating across application, infrastructure, architecture, cloud, and business teams.

The ideal candidate will automate recovery workflows, develop IaC pipelines, and define SLIs/SLOs to meet RTO/RPO targets while ensuring

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field or equivalent work experience.
  • Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations.
  • Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and experience with infrastructure-as-code (e.g., Terraform, Ansible).
  • Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure).
  • Understanding of high-availability design: redundancy, replication, failover, and load balancing.
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams.

Responsibilities

  • Design and code automation that reduces operational toil and replaces manual, error-prone runbooks with orchestrated, auditable failover and recovery workflows.
  • Develop and maintain infrastructure-as-code, scripts, and pipelines (e.g., Python, Bash, PowerShell, Terraform, Ansible) to provision, configure, and validate recovery environments.
  • Build self-healing patterns, health checks, and automated validation that confirm recoverability before a disaster is ever declared.
  • Partner with application and infrastructure teams to assess system architecture for reliability, redundancy, and recoverability against assigned system criticality.
  • Define, measure, and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services.
  • Ensure architecture is selected to meet Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.
  • Plan and execute disaster recovery tests and targeted fault-injection / chaos experiments to proactively expose weaknesses.
  • Improve monitoring, alerting, and observability to detect service degradation early.
  • Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs, aligning technical and business stakeholders toward clear outcomes.
  • Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams.
  • Produce clear, examiner-ready documentation and evidence, ensuring work aligns with KeyBank policies, standards, and regulatory requirements.

Skills

Python
Bash
PowerShell
Terraform
Ansible
GCP
Azure
On-prem infra
Cloud platforms
DevOps

Education

Bachelor's degree or equivalent

Tools

Dynatrace
Prometheus
Grafana
Splunk
Kubernetes
CI/CD

Job description

KeyBank is seeking a Resiliency Engineer to design, build, and improve reliability, availability, and recoverability across on-premises, hybrid, and cloud environments. This hands-on role requires coding skills and a capacity to operate during live recovery exercises, coordinating across application, infrastructure, architecture, cloud, and business teams.

The ideal candidate will automate recovery workflows, develop IaC pipelines, and define SLIs/SLOs to meet RTO/RPO targets while ensuring

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Resiliency Engineer & Automation Lead
Cloud Resiliency Engineer & Automation Lead

KeyBank • City of Albany (NY)

On-site
USD 63,000 - 96,000
Resiliency Engineer: Cloud Reliability & DR (Remote)
Resiliency Engineer: Cloud Reliability & DR (Remote)

KeyBank • Kentucky

On-site
USD 63,000 - 96,000
Site Reliability Engineer — Resiliency & Automation
Site Reliability Engineer — Resiliency & Automation

KeyBank • United States

On-site
USD 63,000 - 96,000
Resiliency Engineer
Resiliency Engineer

KeyBank • City of Albany (NY)

On-site
USD 63,000 - 96,000
Resiliency Engineer
Resiliency Engineer

KeyBank • Brooklyn (OH)

On-site
USD 63,000 - 96,000
Resiliency Engineer
Resiliency Engineer

KeyCorp • City of Albany (NY)

Hybrid
USD 63,000 - 96,000
In-office presence with flexible opts
Competitive base salary
Resiliency Engineer
Resiliency Engineer

KeyBank • Kentucky

On-site
USD 63,000 - 96,000
Resiliency Engineer
Resiliency Engineer

KeyBank • United States

On-site
USD 63,000 - 96,000
Senior Resiliency Architect (Cloud, DR, Observability)
Senior Resiliency Architect (Cloud, DR, Observability)

Huntington Bank • Kentucky

On-site
USD 70,000 - 140,000
Health insurance
Wellness program
Life and disability insurance
+4
Senior Resiliency Architect: Cloud, DR & Reliability
Senior Resiliency Architect: Cloud, DR & Reliability

Huntington National Bank • Chicago (IL)

On-site
USD 70,000 - 140,000
health insurance coverage
wellness program
life and disability insurance
+4