Platform Engineer IV

capgroup

Charlotte (NC)

Hybrid

USD 150,000 - 210,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Matching gifts program
Annual charitable grants
Professional development resources
Flexible work options
Health benefits
Generous time away

Job summary

Capital Group is seeking a Senior Infrastructure & Application Resiliency Engineer to architect and evolve fault-tolerant AWS patterns. You will lead cross-functional resiliency work with product, application, and platform teams, defining reliability strategy and SLOs, and coordinating incident response for high-severity outages.

You will design multi-region architectures, publish reference patterns, mentor engineers, and drive operational excellence across the enterprise.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent cloud experience.
  • At least 10 years in SRE/DevOps/Cloud Architecture with enterprise-scale leadership.
  • Deep expertise in backup, DR, and cyber resilience on AWS and cross-region recovery.
  • Expert in AWS resiliency services, multi-region architectures, and Kubernetes orchestration.
  • Experience defining SLOs/SLIs, and automated self-healing with strict RTO/RPO targets.
  • Hands-on chaos/failure-injection with Chaos Mesh, Gremlin, or AWS FIS.
  • IaC skills (Terraform) and observability with Prometheus/OpenTelemetry/Datadog; AutoSys/Temporal/Kafka a plus.
  • Mentor engineers and communicate complex tech to executives.

Responsibilities

  • Architect, build, and evolve fault-tolerant infrastructure and patterns in AWS.
  • Lead cross-functional resiliency initiatives and define reliability strategy (SLOs, error budgets).
  • Design multi-region, multi-AZ architectures and publish enterprise reference patterns.
  • Coordinate incident response for high-severity outages and perform blameless post-mortems.
  • Mentor senior engineers and elevate organizational reliability practices.

Skills

SRE/DevOps
AWS
Kubernetes
Terraform
Chaos engineering
Observability
Leadership

Education

Bachelor's in CS/Engineering

Tools

Chaos Mesh
Gremlin
AWS FIS
EKS
Datadog

Job description

"I can be myself at work."

You are more than a job title. We want you to feel comfortable doing great work and bringing your best, authentic self to everything you do. We value your talents, traditions, and uniqueness-and we're committed to fostering a strong sense of belonging in a respectful workplace.

We intentionally seek diverse perspectives, experiences, and backgrounds, investing in a culture designed to celebrate differences. We believe that belonging leads to better outcomes and a stronger community of associates united by our mission. At Capital, we live our core values every day: Integrity, Client Focus, Diverse Perspectives, Long-Term Thinking, and Community.

I can influence my income.

You want to feel recognized at work. Your performance will be reviewed annually, and your compensation will be designed to motivate and reward the value that you provide. You'll receive a competitive salary, bonuses and benefits. Your company-funded retirement contribution will factor in salary and variable pay, including bonuses.

I can lead a full life.

You bring unique goals and interests to your job and your life. Whether you're raising a family, you're passionate about where you volunteer, or you want to explore different career paths, we'll give you the resources that can set you up for success. Enjoy generous time-away and health benefits from day one, with the opportunity for flexible work options

Receive 2-for-1 matching gifts for your charitable contributions and the opportunity to secure annual grants for the organizations you love

Access on-demand professional development resources that allow you to hone existing skills and learn new ones

I can succeed as a Senior Infrastructure & Application Resiliency Engineer at Capital Group.

On our Cloud Reliability Engineering team, you'll architect, build, and evolve the fault-tolerant infrastructure and application patterns the enterprise depends on to meet its availability and recovery commitments. Working autonomously and with wide latitude, you'll serve as a strategic technical lead for the most complex, cross-functional resiliency solutions running in AWS-partnering with product, application, and platform leaders to define reliability strategy, from SLOs and error budgets to disaster-r ecovery posture.

You’ll design multi-region, multi-AZ architectures and publish the reference patterns teams adopt across the enterprise; lead chaos experiments and game days that prove out resilience; automate self-healing and orchestrated failover to hit strict RTO/RPO targets; and coordinate incident response for high-severity outages, authoring the blameless post-mortems that permanently fix failures. Along the way, you'll mentor senior and lead engineers and raise the bar for reliability practices org-wide .

I am the person Capital Group is looking for.
  • You have a bachelor's degree in Computer Science , Engineering, or equivalent practical cloud experience
  • You have at least 10 years in Site Reliability Engineering (SRE), DevOps, or Cloud Architecture roles, including demonstrated experience as a technical lead on enterprise-scale reliability initiatives
  • You have deep expertise in enterprise backup, recovery, disaster recovery, and cyber resilience on AWS-AWS Backup, cross-region recovery, immutable backups, recovery orchestration, application-consistent recovery, RTO/RPO planning, recovery testing, and backup governance
  • You are an expert in AWS resiliency services (e.g., Application Recovery Controller and Resiliency Hub), multi-region architectures, modern application architecture, and Kubernetes/container orchestration (e.g., EKS scaling and networking)
  • You have proven experience defining SLOs/SLIs and error budgets and automating self-healing and orchestrated failover to meet strict RTO/RPO targets
  • You have hands-on experience leading chaos/failure-injection experiments with tools such as Chaos Mesh, Gremlin, or AWS Fault Injection Simulator (FIS)
  • You have advanced Infrastructure-as-Code skills (e.g., Terraform for immutable infrastructure) and observability expertise with Prometheus, OpenTelemetry , or Datadog; working knowledge of workflow/streaming platforms such as AutoSys, Temporal.io, and Kafka is a plus
  • You have proven ability to mentor engineers, drive operational excellence, and articulate complex technical challenges and solutions to business and IT executive audiences
I can learn more about Capital Group.

At Capital Group, the success of the people who invest with us depends on the people in whom we invest. That's why we offer a culture, compensation and opportunities that empower our associates to build successful and prosperous c

Get your free, confidential resume review.
or drag and drop your file here.