Principal AWS Resiliency Architect

Insight Global

Austin (TX)

Hybrid

USD 180,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Insight Global is seeking a Principal Engineer to own the resiliency strategy and reference architecture for a large-scale AWS environment in Austin. You will establish availability standards, lead business continuity and disaster recovery initiatives, and drive high-availability architecture across infrastructure, containers, data platforms, and engineering teams.

This role requires 12+ years of engineering, hands-on AWS across multi-account/multi-region environments, and strong

Qualifications

  • 12+ years of engineering experience with senior ownership of resiliency or high-availability architecture.
  • Hands-on AWS expertise across multi-account/multi-region environments, networking, IAM, Route 53, load balancing, and Global Accelerator.
  • Proven ownership of multi-AZ/multi-region architecture, BC/DR strategy, automated failover, and recovery testing.
  • Expertise with AWS Well-Architected Framework and production container platforms (EKS, ECS/Fargate).
  • Data resiliency with Aurora/RDS and at least one of DynamoDB, ElastiCache, or S3 replication.
  • Advanced Terraform/IaC and infrastructure CI/CD, including state management and drift detection.
  • Chaos engineering, observability, SLOs, error budgets, and incident response.
  • Strong cross-functional leadership balancing availability, cost, risk and complexity.

Responsibilities

  • Own AWS resiliency strategy and availability standards across multi-AZ/multi-region architecture.
  • Define patterns for active-active, active-passive, warm-standby, and pilot-light deployments.
  • Lead BC/DR planning with RTO/RPO targets, automated failover, runbooks and disaster testing.
  • Conduct AWS Well-Architected Reviews and drive remediation.
  • Build resiliency across ECS/Fargate, EKS, Aurora, RDS, DynamoDB, ElastiCache, and S3.
  • Automate recovery with IaC, self-healing, drift detection, and fault injection.
  • Establish SLOs, health checks, dependency mapping and observability standards.
  • Lead availability incident response and translate failures into architectural improvements.
  • Collaborate with engineering, product, finance and business leaders to balance reliability, cost, and risk.

Skills

AWS resiliency
Multi-account AWS
BC/DR planning
RTO/RPO
EKS & ECS/Fargate
Aurora/RDS & DynamoDB/ElastiCache/S3
Terraform/IaC
Chaos engineering
Observability & SLOs
Incident response
Cross-functional leadership

Tools

Terraform

Job description

  • 12+ years of engineering experience, including 7+ years owning large-scale resiliency or high-availability architecture.
  • Deep hands-on AWS experience across multi-account, multi-region environments, networking, IAM, Route 53, load balancing, and Global Accelerator.
  • Proven ownership of multi-AZ/multi-region architecture, BC/DR strategy, automated failover, RTO/RPO, and recovery testing.
  • Expertise in the AWS Well-Architected Framework and production container platforms, including EKS and ECS/Fargate.
  • Strong data resiliency experience with Aurora/RDS and at least one of DynamoDB, ElastiCache, or S3 replication.
  • Advanced Terraform/IaC and infrastructure CI/CD experience, including state management, guardrails, and drift detection.
  • Hands-on experience with chaos engineering, observability, SLOs, error budgets, and incident response.
  • Strong cross-functional leadership with the ability to balance availability, cost, risk, and operational complexity.
Nice to Have Skills & Experience
  • Experience with AWS resiliency tools, infrastructure orchestration platforms, and SRE practices.
  • AWS Solutions Architect Professional certification or equivalent expertise.
  • Experience modeling redundancy costs and presenting tradeoffs to business stakeholders.
  • Background integrating AWS environments after a merger or acquisition.
  • Experience supporting regulated or mission-critical environments.
Job Description

Seeking a Principal Engineer to own the resiliency strategy and reference architecture for a large-scale AWS environment. This individual will establish availability standards, lead business continuity and disaster recovery initiatives, and drive high-availability architecture across infrastructure, containers, data platforms, and engineering teams.

Key Responsibilities

- Own AWS resiliency strategy, availability standards, and multi-AZ/multi-region architecture.

- Define appropriate active-active, active-passive, warm-standby, and pilot-light patterns.

- Lead BC/DR planning, including RTO/RPO targets, automated failover, runbooks, DR testing, and game days.

Conduct AWS Well-Architected Reviews and drive remediation efforts.

- Build resiliency across ECS/Fargate, EKS, Aurora, RDS, - DynamoDB, ElastiCache, and S3.

- Automate recovery using IaC, self-healing, drift detection, and fault injection.

- Establish SLOs, error budgets, health checks, dependency mapping, and observability standards.

- Lead availability incident response and convert recurring failures into architectural improvements.

- Partner with engineering, product, finance, and business leaders to balance reliability, cost, and operational complexity.

This is a hybrid position in Austin Texas and pays between $180,000 and $210,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Engineer - Resiliency
Principal Engineer - Resiliency

Insight Global • Austin (TX)

On-site
USD 180,000 - 280,000
Principal AWS Resiliency Architect
Principal AWS Resiliency Architect

Insight Global • Austin (TX)

On-site
USD 180,000 - 280,000
AWS Resiliency Architect - Multi-Region HA Leader
AWS Resiliency Architect - Multi-Region HA Leader

Insight Global • Austin (TX)

Hybrid
USD 180,000 - 210,000
AWS Architect
AWS Architect

Sovereign Technologies, LLC. • Boston (MA)

On-site
USD 160,000 - 210,000
Resiliency Architect
Resiliency Architect

ALLTECH CONSULTING SVC INC • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Senior Software Development Engineer, AWS Resilience, Incident Prevention
Senior Software Development Engineer, AWS Resilience, Incident Prevention

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Software Dev Engineer, AWS Resilience Hub
Software Dev Engineer, AWS Resilience Hub

Amazon • Portland (OR)

On-site
USD 143,700 - 194,400
Health insurance
401(k) matching
Paid time off
+1
AWS Architect & Platform Engineering Lead
AWS Architect & Platform Engineering Lead

Micrologic • Parsippany-Troy Hills (NJ)

On-site
USD 140,000 - 190,000
Software Dev Engineer, AWS Resilience Hub
Software Dev Engineer, AWS Resilience Hub

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
AWS/Java Technical Architect
AWS/Java Technical Architect

Sovereign Technologies, LLC. • Northern (KY)

Hybrid
USD 140,000 - 200,000