Resiliency Engineer

KeyBank

City of Albany (NY)

On-site

USD 63,000 - 96,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

KeyBank is seeking a Resiliency Engineer to design, build, and continuously improve the reliability of technology platforms across on-prem, hybrid, and cloud environments (GCP/Azure). This hands-on role emphasizes automation, distributed-system thinking, and collaboration with application, infrastructure, and cloud teams to meet recovery objectives.

The ideal candidate writes code, builds automation, and can participate in live recovery exercises while aligning with regulatory requirements and

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field or equivalent work experience.
  • Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations.
  • Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and IaC tooling (Terraform, Ansible).
  • Working knowledge of on-premises infrastructure and public cloud platforms (GCP and/or Azure).
  • Understanding of high-availability design: redundancy, replication, failover, and load balancing.
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills.

Responsibilities

  • Design and code automation that reduces toil and enables automated recovery workflows.
  • Develop and maintain infrastructure-as-code, scripts, and pipelines to provision and validate recovery environments.
  • Build self-healing patterns and automated validation to prove recoverability.
  • Partner with teams to assess architecture for reliability, redundancy, and recoverability.
  • Plan disaster recovery tests and chaos experiments to reveal weaknesses.
  • Improve monitoring, alerting, and observability for early degradation detection.
  • Facilitate resiliency reviews and cross-team recovery walkthroughs; produce clear documentation.

Skills

Automation coding
Python
Bash
PowerShell
Terraform
Ansible
Reliability engineering
DevOps
Cloud platforms
On-prem & cloud
SRE concepts

Education

Bachelor's degree in CS/IT/Engineering or equivalent

Tools

Terraform
Ansible
Kubernetes/GKE
Dynatrace
Prometheus
Grafana
Splunk
Cloud platforms (GCP/Azure)

Job description

Location: 555 Patroon Creek Boulevard, Albany New York

Position Summary

The Resiliency Engineer designs, builds, and continuously improves the reliability, availability, and recoverability of KeyBank's technology platforms across on-premises, hybrid, and cloud (GCP/Azure) environments. Applying software engineering discipline to operations, this role engineers systems to meet defined recovery objectives, automates recovery and validation, and proves that critical services can withstand and recover from failure.

This is a hands-on engineering role for a builder who can write automation, reason about distributed-system failure modes, and facilitate across application, infrastructure, architecture, cloud, and line-of-business teams. The ideal candidate is equally comfortable in code, in a design review, and in a command center during a live recovery exercise.

Key Responsibilities
Automation & Engineering
  • Design and code automation that reduces operational toil and replaces manual, error-prone runbooks with orchestrated, auditable failover and recovery workflows.
  • Develop and maintain infrastructure-as-code, scripts, and pipelines (e.g., Python, Bash, PowerShell, Terraform, Ansible) to provision, configure, and validate recovery environments.
  • Build self-healing patterns, health checks, and automated validation that confirm recoverability before a disaster is ever declared.
Reliability & Resiliency Design
  • Partner with application and infrastructure teams to assess system architecture for reliability, redundancy, and recoverability against assigned system criticality.
  • Define, measure, and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services.
  • Ensure architecture is selected to meet Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.
Testing, Chaos & Validation
  • Plan and execute disaster recovery tests and targeted fault-injection / chaos experiments (e.g., zonal failure, load-balancer, regional failover) to proactively expose weaknesses.
  • Improve monitoring, alerting, and observability to detect service degradation early.
Facilitation & Partnership
  • Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs, aligning technical and business stakeholders toward clear outcomes.
  • Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams.
  • Produce clear, examiner-ready documentation and evidence, ensuring work aligns with KeyBank policies, standards, and regulatory requirements.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field-or equivalent work experience.
  • Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations.
  • Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and experience with infrastructure-as-code (e.g., Terraform, Ansible).
  • Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure).
  • Understanding of high-availability design: redundancy, replication, failover, and load balancing.
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams.
Preferred Qualifications
  • Experience with site reliability engineering principles and service-level management (SLIs, SLOs, error budgets).
  • Experience with disaster recovery planning, resiliency testing, or chaos / fault-injection engineering (e.g., Google FIT, Gremlin).
  • Familiarity with containers and orchestration (Kubernetes/GKE), CI/CD, and observability tooling (e.g., Dynatrace, Prometheus, Grafana, Splunk).
  • Experience with ServiceNow (ITOM / Business Continuity Management) or comparable orchestration platforms.
  • Experience in a regulated industry or large, complex enterprise environment; familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301.
  • Relevant certifications (e.g., cloud architect/engineer, Kubernetes, Linux, ITIL).
Compensation and Benefits

This position is eligible to earn a base salary in the range of $63,000.00 - $96,000.00 annually. Placement within the pay range may differ based upon various factors, including but not limited to skills, experience and geographic location. Compensation for this role also includes eligibility for incentive compensation which may include production, commission, and/or discretionary incentives.

Key has implemented an approach to employee workspaces which prioritizes in-office presence, while providing flexible options in circumstances where roles can be performed effectively in a mobile environment.

KeyCorp is an Equal Opportunity Employer committed to sustaining an inclusive culture. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, genetic information, pregnancy, disability, veteran status or any other characteristic protected by law.

Qualified individuals with disabilities or disabled veterans who are unable or limited in their ability to apply on this site may request reasonable accommodations by emailing HR_Compliance@keybank.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Resiliency Engineer
Resiliency Engineer

KeyBank • Brooklyn (OH)

On-site
USD 63,000 - 96,000
Resiliency Engineer
Resiliency Engineer

KeyBank • United States

On-site
USD 63,000 - 96,000
Resiliency Engineer
Resiliency Engineer

KeyCorp • City of Albany (NY)

Hybrid
USD 63,000 - 96,000
In-office presence with flexible opts
Competitive base salary
Resiliency Engineer
Resiliency Engineer

KeyBank • Kentucky

On-site
USD 63,000 - 96,000
Resiliency Engineer: Cloud Reliability & DR (Remote)
Resiliency Engineer: Cloud Reliability & DR (Remote)

KeyBank • Kentucky

On-site
USD 63,000 - 96,000
Cloud Resiliency Engineer - Automation & DR Leader
Cloud Resiliency Engineer - Automation & DR Leader

KeyBank • Brooklyn (OH)

On-site
USD 63,000 - 96,000
Site Reliability Engineer — Resiliency & Automation
Site Reliability Engineer — Resiliency & Automation

KeyBank • United States

On-site
USD 63,000 - 96,000
Cloud Resiliency Engineer & Automation Lead
Cloud Resiliency Engineer & Automation Lead

KeyBank • City of Albany (NY)

On-site
USD 63,000 - 96,000
Senior Resiliency Engineer-2
Senior Resiliency Engineer-2

Huntington Bancshares, Inc. • Columbus (OH)

On-site
USD 93,000 - 189,000
Health insurance
Wellness program
Life insurance
+4
Cyber Defense Incident Response Lead
Cyber Defense Incident Response Lead

KeyBank • Brooklyn (OH)

On-site
USD 96,000 - 181,000