Site Reliability Engineer

Jobtailor

United States

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a Cloud Reliability Engineer to maintain SLOs for cloud-hosted systems and design automation across AWS and Azure. You will collaborate with development teams to improve reliability, observability, and release velocity while participating in on-call rotations and postmortems.

The role emphasizes IaC using Terraform, containerized workloads with Docker and Kubernetes, and robust CI/CD pipelines.

Qualifications

  • Bachelor's degree or equivalent experience in Computer Science or related field.
  • Proficiency in scripting languages (Shell, Python).
  • Experience maintaining Infrastructure as Code (IaC) using Terraform (or AWS CloudFormation, Azure ARM templates).
  • Hands-on experience with AWS services (EC2, S3, RDS, IAM, VPC, Lambda) and familiarity with Azure equivalents.
  • Knowledge of Docker and Kubernetes (EKS or AKS preferred).
  • Familiarity with observability tools like Datadog, AWS CloudWatch, or Azure Monitor.
  • Implement and maintain CI/CD pipelines (AWS CodePipeline, Azure DevOps, or similar).
  • Incident response and blameless postmortems.
  • Proficient in Git workflows.
  • 2-5 years of proven experience.

Responsibilities

  • Maintain SLOs for cloud-hosted systems with automation, reliability, and observability across AWS and Azure.
  • Design and implement automation for AWS and Azure to scale systems sustainably and enable rapid recovery.
  • Partner with development teams to improve reliability and release velocity using cloud-native tools and guidelines.
  • Participate in on-call rotations, incident response, postmortems, and root cause analysis.
  • Advocate for strong engineering practices for building, deploying, and running scalable services across AWS and Azure.
  • Enable cloud migration by conducting architectural reviews and readiness testing, and configuring observability dashboards.
  • Drive continuous learning and development in multi-cloud technologies.

Skills

Shell scripting
Python
Git workflows

Education

Bachelor's degree or equivalent in CS or related field

Tools

Terraform
AWS CloudFormation
Azure ARM templates
Datadog
AWS CloudWatch
Azure Monitor
AWS CodePipeline
Azure DevOps
Docker
Kubernetes (EKS/AKS)

Job description

  • Maintain Service Level Objectives (SLOs) for cloud-hosted systems, continuously measuring and improving availability, latency, and overall system health.
  • Design and implement automation for AWS and Azure environments to scale systems sustainably, prevent service issues, and enable rapid recovery.
  • Partner with development teams to improve reliability, observability, and release velocity using cloud-native tools and guidelines.
  • Participate in on-call rotations, incident response, postmortems, and root cause analysis.
  • Advocate for strong engineering practices for building, deploying, and running scalable, reliable services across AWS and Azure.
  • Enable cloud migration by performing architectural reviews, operational readiness testing, and configuring observability dashboards (e.g., Datadog integrated with AWS CloudWatch and Azure Monitor).
  • Drive continuous learning and development in multi-cloud technologies.
Requirements
  • Bachelor's degree or equivalent experience in Computer Science or related technical field.
  • Proficiency in scripting languages (Shell, Python).
  • Experience maintaining Infrastructure as Code (IaC) using Terraform (or AWS CloudFormation, Azure ARM templates).
  • Hands‑on experience with AWS services (EC2, S3, RDS, IAM, VPC, Lambda) and familiarity with Azure equivalents.
  • Knowledge of Docker and Kubernetes (EKS or AKS preferred).
  • Familiarity with observability tools like Datadog, AWS CloudWatch, or Azure Monitor.
  • Implement and maintain CI/CD pipelines (AWS CodePipeline, Azure DevOps, or similar).
  • Incident response and blameless postmortems.
  • Proficient in Git workflows.
  • 2-5 years of proven experience.
Core Competencies

Demonstrates expertise in maintaining Service Level Objectives (SLOs) for cloud-hosted systems, with a strong focus on automation, reliability, and observability across AWS and Azure environments. Proficient in Infrastructure as Code (IaC) practices and incident response methodologies to ensure scalable and reliable service delivery.

Highest-signal resume keywords
  • AWS Services (EC2, S3, RDS, IAM, VPC, Lambda)
  • Infrastructure as Code (IaC) Using Terraform
  • Scripting Languages (Shell, Python)
  • CI/CD Pipeline Implementation (AWS CodePipeline, Azure DevOps)
  • Docker and Kubernetes (EKS or AKS)
ATS Optimization Keywords
Hard Skills
  • AWS Services
  • Azure Services
  • Infrastructure as Code (IaC)
  • Scripting Languages
  • CI/CD Pipelines
  • Docker
  • Kubernetes
  • Git Workflows
  • Observability Tools
  • Incident Response
Soft Skills
  • Collaboration
  • Problem-Solving
  • Continuous Learning
Industry Keywords
  • Service Level Objectives (SLOs)
  • Cloud Migration
  • Operational Readiness Testing
  • Postmortems
  • Root Cause Analysis
Tools & Technologies
  • Terraform
  • AWS CloudFormation
  • Azure ARM Templates
  • Datadog
  • AWS CloudWatch
  • Azure Monitor
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Jobtailor • California (MO)

On-site
USD 85,000 - 120,000
Software Engineer – AWS, Python, DevOps
Software Engineer – AWS, Python, DevOps

Jobtailor • Connecticut

On-site
USD 110,000 - 150,000
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Jobtailor • Bethesda (MD)

On-site
USD 120,000 - 180,000
Senior Application Development Advisor
Senior Application Development Advisor

Jobtailor • Colorado

Hybrid
USD 130,000 - 170,000
Senior Engineer, IT Infrastructure Engineering – Data Center
Senior Engineer, IT Infrastructure Engineering – Data Center

Jobtailor • Atlanta (GA)

Hybrid
USD 140,000 - 180,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Cloud Infrastructure DevOps Engineer
Cloud Infrastructure DevOps Engineer

Jobtailor • Austin (TX)

On-site
USD 140,000 - 190,000
Cloud Engineer
Cloud Engineer

Jobtailor • Idaho Falls (ID)

On-site
USD 120,000 - 150,000
Software Engineering Lead
Software Engineering Lead

Jobtailor • Colorado

Hybrid
USD 170,000 - 210,000
Platform Automation Engineer
Platform Automation Engineer

Jobtailor • Maryland

On-site
USD 120,000 - 180,000