Sr. Staff Cloud Resilience Engineer - Security

UKG

Atlanta (GA)

Hybrid

USD 140,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

UKG seeks a senior resilience architect to advise distributed teams on high availability and redundancy for SaaS applications on Google Cloud, with oversight for AWS and Azure workloads.

You will review IaC implementations, design fast-failover strategies, and ensure self-healing deployments on GKE and Compute Engine while collaborating with DevOps, SRE, and Product Engineering for resilience and failover validation.

Qualifications

  • Proven experience designing HA architectures for SaaS workloads.
  • Strong background in IaC review and guidance for engineering teams.
  • Hands-on engineer with implementation insights to guide others.

Responsibilities

  • Advise on resilience practices for distributed product and enterprise teams.
  • Design fast-failover solutions and highly available infrastructure on GCP; oversee AWS/Azure workloads.
  • Lead chaos testing and validate failover across tier-0 systems.

Skills

HA & DR
Chaos testing
SRE collaboration
Cross-functional leadership
Resilience engineering

Tools

GKE
Compute Engine
CloudSQL
DNS & Networking
SIEM

Job description

Architectural Leadership & Consulting

Act as the primary resilience advisor to multiple distributed product and enterprise teams, guiding them on best practices for building high availability (HA) and redundancy into their SaaS applications.

Resilient Cloud Design

Design and recommend fast-failover solutions and highly available infrastructure primarily on Google Cloud Platform (GCP), while also providing oversight for workloads in Azure and AWS.

Infrastructure Validation

Leverage your strong background in Infrastructure as Code (IaC) to review, validate, and guide the implementation efforts of engineering teams.

Container & Compute Resilience

Design redundancy strategies for workloads running on Google Kubernetes Engine (GKE) and virtual machines, ensuring self‑healing deployments.

Cross‑Functional Collaboration

Partner closely with DevOps, SRE, and Product Engineering teams to champion resilience engineering principles, chaos testing, and failover validations across tier‑0 mission‑critical systems.

Cloud Platform Expertise

Deep, practical technical knowledge of Google Cloud Platform (GCP) core services, specifically GKE, Compute Engine, and CloudSQL. Familiarity with AWS and Azure is highly desirable.

Technical Practitioner Background

Proven past experience as a hands‑on engineer who has deployed complex infrastructure. You should understand the implementation details well enough to effectively guide engineering teams.

High Availability Architecture

Demonstrated success in architecting active‑active or active‑passive fast‑failover mechanisms for high‑volume, data‑intensive SaaS applications.

Database Resilience

Strong understanding of database clustering, replication, and migration strategies (especially migrating legacy RDBMS like MS SQL Server to cloud‑native solutions like CloudSQL).

Advisory Skills

Excellent communication and consulting skills, with the ability to influence technical teams, explain complex architectural concepts, and foster a culture of resilience without having direct reporting authority over the engineering teams.

Resilient Network Services

Practical design experience managing high‑availability network topologies, including load balancing, DNS & name resolution, firewalls/gateways, identity/authentication systems, and centralized logging/SIEM.

User Session Management

Deep understanding of user session replication, session state persistence, and failover routing strategies in high‑traffic, multi‑region application architectures.

Broad HA Domain Exposure

Familiarity assessing or designing resilience across a comprehensive range of critical SaaS failure domains, such as API gateways, caching layers, messaging/queuing systems, and CI/CD pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Resilience Engineer: HA, Chaos Testing & SRE
Senior Cloud Resilience Engineer: HA, Chaos Testing & SRE

UKG • Atlanta (GA)

Hybrid
USD 140,000 - 190,000
Senior Cloud Resilience Architect (GCP)
Senior Cloud Resilience Architect (GCP)

UKG • Lowell (MA)

On-site
USD 129,000 - 187,000
Cloud Resilience Architect: GCP & Multi-Cloud HA
Cloud Resilience Architect: GCP & Multi-Cloud HA

UKG • Seattle (WA)

On-site
USD 129,000 - 187,000
Resiliency Architect
Resiliency Architect

ALLTECH CONSULTING SVC INC • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Sr. Staff Cloud Resilience Engineer - Security
Sr. Staff Cloud Resilience Engineer - Security

UKG • Seattle (WA)

On-site
USD 129,000 - 187,000
Sr. Staff Cloud Resilience Engineer - Security
Sr. Staff Cloud Resilience Engineer - Security

UKG • San Francisco (CA)

On-site
USD 129,000 - 187,000
Sr. Staff Cloud Resilience Engineer - Security
Sr. Staff Cloud Resilience Engineer - Security

UKG • Town of Florida (NY)

On-site
USD 129,000 - 187,000
Performance-based bonus plan
Restricted stock unit awards
Comprehensive benefits and rewards
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
Cloud Resilience Architect: GCP & Multi-Cloud HA
Cloud Resilience Architect: GCP & Multi-Cloud HA

UKG • Town of Florida (NY)

On-site
USD 129,000 - 187,000
Performance-based bonus plan
Restricted stock unit awards
Comprehensive benefits and rewards
Sr. Staff Cloud Resilience Engineer - Security
Sr. Staff Cloud Resilience Engineer - Security

UKG • Lowell (MA)

On-site
USD 129,000 - 187,000