Remote SRE: AWS, Kubernetes & Observability

Cloudbeds

United States

Remote

USD 120,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote First
PTO
Home office stipend
Cloudbeds University
Upskilling opportunities

Job summary

Cloudbeds seeks a Site Reliability Engineer to safeguard the reliability and performance of our hospitality platform. You will architect scalable AWS cloud solutions and maintain Kubernetes (EKS) clusters, while supporting CI/CD with ArgoCD and GitHub Actions, and automating deployments with Terraform.

You will develop and improve product observability with Grafana, Prometheus, DataDog, and CloudWatch, and participate in incident management and RCA to minimize downtime.

Qualifications

  • 5+ years as a DevOps or SRE in AWS environments.
  • Strong Kubernetes (EKS) experience and Helm charts.
  • CI/CD design with ArgoCD and GitHub Actions.
  • Infrastructure-as-code with Terraform.
  • Observability with Grafana/Prometheus/DataDog/CloudWatch.

Responsibilities

  • Design and run reliable, scalable AWS architecture.
  • Maintain and support large Kubernetes (EKS) clusters.
  • Support CI/CD with ArgoCD and GitHub Actions.
  • Automate deployments with Terraform IaC.
  • Improve observability and monitoring across platforms.
  • Lead incident management and RCA activities.
  • Collaborate with dev teams on reliability best practices.

Skills

AWS
Kubernetes
CI/CD
Observability
Incident Management
Networking
English communication

Education

Bachelor’s degree in CS or equivalent

Tools

EKS
ArgoCD
GitHub Actions
Terraform
Grafana/Prometheus
Datadog
CloudWatch
Nginx/Ingress

Job description

Cloudbeds seeks a Site Reliability Engineer to safeguard the reliability and performance of our hospitality platform. You will architect scalable AWS cloud solutions and maintain Kubernetes (EKS) clusters, while supporting CI/CD with ArgoCD and GitHub Actions, and automating deployments with Terraform.

You will develop and improve product observability with Grafana, Prometheus, DataDog, and CloudWatch, and participate in incident management and RCA to minimize downtime.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE - AWS, Terraform & Kubernetes
Remote SRE - AWS, Terraform & Kubernetes

GiveCampus • United States

Remote
USD 120,000 - 180,000
Remote-friendly
Office in Washington, DC
Remote SRE: Scale Reliability with Cloud Automation
Remote SRE: Scale Reliability with Cloud Automation

Glassbox • New York (NY)

On-site
USD 120,000 - 150,000
Remote Senior SRE — Observability & Cloud
Remote Senior SRE — Observability & Cloud

Cribl • Boise (ID)

Remote
USD 142,000 - 195,000
Health
Dental
Vision
+7
Remote Senior SRE - Observability & Cloud
Remote Senior SRE - Observability & Cloud

Cribl • Frankfort (KY)

Remote
USD 142,000 - 195,000
Health insurance
Dental
Vision
+7
Senior SRE - Kubernetes, Observability & Automation (Remote)
Senior SRE - Kubernetes, Observability & Automation (Remote)

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote work
Annual company events
Health & wellbeing
+2
Remote DevOps Engineer: AWS, Kubernetes & CI/CD
Remote DevOps Engineer: AWS, Kubernetes & CI/CD

Cloudbeds • Atlanta (GA)

On-site
USD 90,000 - 120,000
Home office stipend based on country
Fully Paid Parental Leave
Monthly Wellness Fridays
+2
Senior Site Reliability Engineer - Cloud, Kubernetes & Automation
Senior Site Reliability Engineer - Cloud, Kubernetes & Automation

Socure • Town of Concord (NY)

On-site
USD 140,000 - 210,000
Remote Senior SRE: AWS, Kubernetes & Terraform
Remote Senior SRE: AWS, Kubernetes & Terraform

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior SRE: Cloud Reliability & AI-Driven Incident Triage
Senior SRE: Cloud Reliability & AI-Driven Incident Triage

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Senior Cloud SRE — AI-Driven Uptime & Kubernetes
Senior Cloud SRE — AI-Driven Uptime & Kubernetes

Clearwater Analytics, LLC • Boise (ID)

On-site
USD 130,000 - 170,000