Remote Azure SRE – Observability, CI/CD & Resilience

System Automation Corporation

United States

On-site

USD 120,000 - 140,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

System Automation Corporation seeks a mid-level Site Reliability Engineer to build and operate our Azure-based cloud platform. You will define observability standards, automate toil, and help the team ship changes safely via CI/CD, all while sharing an on-call rotation for incident response.

Reporting to the Manager of Platform Operations and Compliance, this role is 100% remote and eligible to work in the U.S.; compensation includes a base salary plus potential commission and profit sharing.

Qualifications

  • 3+ years in IT Operations, DevOps, or SRE.
  • Hands-on Azure experience in production.
  • Experience with Terraform and/or Bicep.
  • Scripting or programming in Python or TypeScript.
  • Experience with REST and GraphQL APIs.
  • Experience defining and tracking KPIs/SLOs for web apps.
  • Comfortable participating in an on-call rotation.

Responsibilities

  • Build, operate, and scale production systems in Microsoft Azure to meet availability and performance targets.
  • Triage and resolve production incidents; contribute to blameless postmortems.
  • Define and track SLOs/SLIs and error budgets with engineering and product teams.
  • Design observability tooling (metrics, logs, tracing, dashboards) and refine alerting.
  • Build and maintain CI/CD pipelines enabling safe, frequent deployments.
  • Provision and manage infrastructure as code (Bicep/Terraform) with change control.
  • Ensure security and compliance requirements in collaboration with the compliance team.
  • Document designs, runbooks, diagrams, and stay current with Azure capabilities.

Skills

Azure
CI/CD
SRE
Scripting
APIs
On-call
Observability

Tools

Terraform
Bicep
Node.js
Power Apps

Job description

System Automation Corporation seeks a mid-level Site Reliability Engineer to build and operate our Azure-based cloud platform. You will define observability standards, automate toil, and help the team ship changes safely via CI/CD, all while sharing an on-call rotation for incident response.

Reporting to the Manager of Platform Operations and Compliance, this role is 100% remote and eligible to work in the U.S.; compensation includes a base salary plus potential commission and profit sharing.

Get your free, confidential resume review.
or drag and drop your file here.