Cloud SRE & Resiliency Engineer - Observability & Chaos

United States Digital Space LLC

United States

Remote

USD 140,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Attractive remuneration package and "_

Job summary

United States Digital Space LLC seeks an experienced Site Reliability Engineer to drive observability, resiliency, and cloud migration initiatives across AWS. You will mentor teams, lead incident response, and champion best practices in Kubernetes, IaC, CI/CD, and automation.

Join a team focused on cloud resiliency, on-call readiness, and continuous learning with international training opportunities.

Qualifications

  • 5+ years of cloud services experience with at least 3 years on AWS.
  • 3+ years in SRE or similar role; strong incident management knowledge.
  • Experience with monitoring, logging, and alerting tools; strong troubleshooting.
  • Advanced knowledge of CI/CD, IaC (Terraform, CloudFormation) and Linux.
  • Ability to mentor others and communicate across teams.

Responsibilities

  • Advise and enforce the resiliency pillar of the Well Architected Framework.
  • Conduct Chaos Engineering experiments and related exercises.
  • Design and improve observability across cloud infrastructure.
  • Lead on-call rotations and incident response.
  • Mentor colleagues and drive culture change toward reliability.
  • Coordinate with IT teams on observability needs and standards.
  • Ensure deployments comply with company standards and best practices.

Skills

Cloud services
AWS cloud
SRE
Kubernetes
CI/CD
Infrastructure as Code
Terraform
Git
Linux
Scripting

Education

BSc/MSc in Computer Science or related field

Tools

Terraform (HCL)
AWS CloudFormation
Kubernetes
Git

Job description

United States Digital Space LLC seeks an experienced Site Reliability Engineer to drive observability, resiliency, and cloud migration initiatives across AWS. You will mentor teams, lead incident response, and champion best practices in Kubernetes, IaC, CI/CD, and automation.

Join a team focused on cloud resiliency, on-call readiness, and continuous learning with international training opportunities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Remote SRE: Observability & Cloud Resilience
Remote SRE: Observability & Cloud Resilience

Skyward • Rockville (MD)

On-site
USD 110,000 - 170,000
Medical, dental, vision insurance
401K with 4% employer contribution
Company provided laptop
+2
Senior Cloud SRE: AWS, Serverless & Incident Response
Senior Cloud SRE: AWS, Serverless & Incident Response

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Senior AWS SRE Consultant – Cloud Observability & DR
Senior AWS SRE Consultant – Cloud Observability & DR

Vertical Relevance • United States

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer – Observability & Cloud
Senior Site Reliability Engineer – Observability & Cloud

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000
Remote Cloud SRE: Observability, CI/CD & Resilience
Remote Cloud SRE: Observability, CI/CD & Resilience

01105 Softpro, LLC • Raleigh (NC)

On-site
USD 110,000 - 180,000
CloudDevs: Senior Site Reliability Engineer (SRE)
CloudDevs: Senior Site Reliability Engineer (SRE)

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Cloud Infrastructure Engineer — OCI & Multi-Cloud
Cloud Infrastructure Engineer — OCI & Multi-Cloud

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 200,000
Competitive pay
Equity
Unlimited PTO
+2
Senior Site Reliability Engineer – Cloud, Kubernetes & IaC
Senior Site Reliability Engineer – Cloud, Kubernetes & IaC

Comcast Corporation • Reston (VA), Northern (KY)

Hybrid
USD 98,000 - 164,000
Senior Cloud SRE: Scale, Automation & Observability
Senior Cloud SRE: Scale, Automation & Observability

Carrier Global Corporation • Town of Florida (NY)

On-site
USD 96,000 - 192,000
Health Care Benefits
Retirement Benefits
Paid time off