Remote SRE: Observability & Cloud Resilience

Skyward

Rockville (MD)

On-site

USD 110,000 - 170,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401K with 4% employer contribution
Company provided laptop
Paid time off and holidays
Professional development budget

Job summary

Skyward in Rockville, MD is seeking a Site Reliability Engineer to join a mission-driven team modernizing CMS data platforms with AI-driven solutions.

You will manage AWS environments, build observability, and implement IaC with Terraform/Ansible, Jenkins, and Docker, ensuring reliability and cost efficiency. We value collaboration, continuous learning, and a flexible work approach with remote options.

Qualifications

  • A bachelor’s degree in computer science, engineering, or a related field (or equivalent hands-on experience).
  • 3-5 years of experience in site reliability, systems, or cloud engineering, with meaningful time spent in AWS environments.
  • Solid working knowledge of core AWS services, architecture, and best practices.
  • Hands-on experience with infrastructure-as-code tools (Terraform, Ansible, or CloudFormation).
  • A good understanding of CI/CD pipelines and automation tools (Jenkins, GitLab CI, or similar).
  • Comfort scripting and automating in Python.
  • Familiarity with monitoring and observability tooling (CloudWatch, New Relic, Splunk, or comparable).

Responsibilities

  • Join CMS team as it merges and modernizes enterprise knowledge and data systems into a single AI-driven platform.
  • Operate and tune AWS environments to meet infrastructure and application availability SLAs, even during transition.
  • Build observability with dashboards and alerts; establish performance baselines to spot degradation early.
  • Write infrastructure-as-code and support CI/CD pipelines and Docker workloads for repeatable deployments.
  • Define and track SLIs/SLOs, and produce performance/load/bottleneck reports for smarter decisions.
  • Optimize for performance, security, and cost using AWS Trusted Advisor.
  • Support security/compliance modernization and move toward continuous ATO within RMF-boundary.
  • Strengthen resilience with disaster recovery and COOP planning.
  • Own incidents end to end; drive blameless post-mortems and preventative fixes.

Skills

Strong problem-solving
Clear communication
Adaptability

Education

Bachelor's degree in CS/Engineering/related field

Tools

Terraform
Ansible
CloudFormation
Jenkins
GitLab CI
Python
Docker
CloudWatch
New Relic
Splunk

Job description

Skyward in Rockville, MD is seeking a Site Reliability Engineer to join a mission-driven team modernizing CMS data platforms with AI-driven solutions.

You will manage AWS environments, build observability, and implement IaC with Terraform/Ansible, Jenkins, and Docker, ensuring reliability and cost efficiency. We value collaboration, continuous learning, and a flexible work approach with remote options.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Site Reliability Engineer – AI-Driven Cloud Platform
Remote Site Reliability Engineer – AI-Driven Cloud Platform

Skyward IT Solutions, LLC • Rockville (MD)

Hybrid
USD 112,000 - 150,000
Medical, dental, vision insurance (fully paid for employees)
401(k) with 4% employer contribution
Up to 4 weeks of paid parental leave
+1
Remote SRE — Scale, Resilience & Observability
Remote SRE — Scale, Resilience & Observability

Bright Vision Technologies • United States

On-site
USD 100,000 - 150,000
Competitive base salary
Health benefits
Long-term stability
Remote Senior SRE: Resilient Infrastructure & Observability
Remote Senior SRE: Resilient Infrastructure & Observability

Wikimedia Foundation • San Francisco (CA)

On-site
USD 116,000 - 182,000
Competitive salary
Remote-first environment
Inclusive workplace
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)

Motion Recruitment • Chicago (IL)

On-site
USD 140,000 - 190,000
Cloud SRE & Resiliency Engineer - Observability & Chaos
Cloud SRE & Resiliency Engineer - Observability & Chaos

United States Digital Space LLC • United States

Remote
USD 140,000 - 180,000
Attractive remuneration package and "_
Remote Cloud SRE: Observability, CI/CD & Resilience
Remote Cloud SRE: Observability, CI/CD & Resilience

01105 Softpro, LLC • Raleigh (NC)

On-site
USD 110,000 - 180,000
Remote SRE / Platform Engineer — Build Reliable Cloud Platforms
Remote SRE / Platform Engineer — Build Reliable Cloud Platforms

WinTrio LLC • United States

On-site
USD 120,000 - 160,000
Healthcare
401(k)
Annual bonus
+2
Remote SRE Manager: Cloud Reliability & DevOps Lead
Remote SRE Manager: Cloud Reliability & DevOps Lead

Deepwatch • Tampa (FL)

Hybrid
USD 178,000 - 213,000
Medical, dental, vision insurance
Flexible Time Off
Professional development benefits
+1
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Remote Cloud SRE: Build Resilient Health-Tech Infra
Remote Cloud SRE: Build Resilient Health-Tech Infra

UnitedHealth Group • Eden Prairie (MN)

Remote
Confidential
Remote work flexibility
Comprehensive benefits package
Equity stock purchase
+1