Senior Cloud SRE - Observability & Incident Response

Delinea

United States

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Healthcare and other benefits
Bonus program
Pension/retirement matching
Paid time off

Job summary

Delinea is seeking a Senior Site Reliability Engineer to join our Cloud Engineering team. You will own the availability and performance of several production SaaS services running in Azure and AWS, including our FedRAMP High environment.

This hands‑on role emphasizes measuring reliability, tuning detection, responding to incidents, and removing manual work that keeps engineers awake at night. Ideal candidates bring 8+ years in SRE/DevOps, Azure expertise across AKS, App Service, Azure SQL, and a

Qualifications

  • 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.
  • Hands‑on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage.
  • Production experience with an observability platform such as Datadog: metrics, logs, APM, dashboards, and monitor design.
  • Ownership of SLIs, SLOs, and error budgets for services you supported.
  • Incident response experience: you have run a bridge, made calls under pressure, and written the postmortem afterward.
  • Kubernetes in production, including ingress, deployments, resource limits, and troubleshooting failing workloads.
  • Infrastructure as code with Terraform, plus CI/CD pipeline creation and troubleshooting (Azure DevOps preferred).
  • Scripting in PowerShell and Python, and fluency with YAML and JSON.
  • Strong networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet‑level troubleshooting.
  • Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments.

Responsibilities

  • Own reliability for a set of production SaaS services end to end: availability, performance, and capacity.
  • Define service‑level indicators, set service‑level objectives and error budgets, and use them to prioritize reliability work with engineering and product teams.
  • Build and tune monitoring in Datadog and Azure Monitor, including threshold, composite, and anomaly detection monitors, synthetic checks, dashboards, and alert routing that shorten time to detect and cut alert noise.
  • Automate response. Wire monitors into remediation workflows for known failure conditions and replace manual runbook steps with code.
  • Participate in an on‑call rotation, lead incident response for high‑severity events, and coordinate resolution across support, engineering, and product.
  • Write post‑incident reviews and customer‑facing root cause analyses, then drive the preventive actions to completion rather than filing and forgetting them.
  • Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines.
  • Administer the web application firewall: rule tuning, rate limiting, false positive triage, and coordination with the security team.
  • Own the cost of our observability platform: Datadog index and retention spend, custom metric and APM volume, log ingestion rates, and S3 and blob storage for archived logs.
  • Operate within our FedRAMP High environment following established change control and continuous monitoring processes, and contribute to the runbooks and standard operating procedures that support it.
  • Improve how the team runs on‑call: rotation design, escalation paths, alert quality, runbook coverage, and handoff between regions.
  • Partner across Support, Security, Product, and Development to make sure new services ship with monitoring, runbooks, and SLOs in place.

Skills

SRE experience
Azure expert
Datadog observability
SLI/SLO ownership
Incident response
Kubernetes in prod
Terraform IaC
Azure DevOps
PowerShell & Python
Networking fundamentals

Tools

Datadog
Azure Monitor
AKS
Terraform
Azure DevOps
Jenkins
SaltStack
Consul
ELK stack
CloudWatch Logs Insights

Job description

Delinea is seeking a Senior Site Reliability Engineer to join our Cloud Engineering team. You will own the availability and performance of several production SaaS services running in Azure and AWS, including our FedRAMP High environment.

This hands‑on role emphasizes measuring reliability, tuning detection, responding to incidents, and removing manual work that keeps engineers awake at night. Ideal candidates bring 8+ years in SRE/DevOps, Azure expertise across AKS, App Service, Azure SQL, and a

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud SRE: Reliability, Automation & Observability
Senior Cloud SRE: Reliability, Automation & Observability

Delinea • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive salary
Bonus program
Healthcare benefits
+2
Hands-On SRE Manager: Incidents, Observability & Cloud
Hands-On SRE Manager: Incidents, Observability & Cloud

Delinea • Northern (KY)

Hybrid
USD 150,000 - 190,000
Health insurance
Bonus program
Retirement matching
SRE Manager — Incident Command & Platform Reliability
SRE Manager — Incident Command & Platform Reliability

Delinea • United States

On-site
USD 150,000 - 230,000
Healthcare
Pension matching
Life insurance
+3
Senior Cloud SRE: AWS, Serverless & Incident Response
Senior Cloud SRE: AWS, Serverless & Incident Response

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Azure Cloud SRE Lead — Migration, CI/CD & Observability
Azure Cloud SRE Lead — Migration, CI/CD & Observability

Systems Technology Group, Inc. (STG) • Dearborn (MI)

On-site
USD 100,000 - 130,000
Senior SRE Leader: Cloud, Reliability & Scale
Senior SRE Leader: Cloud, Reliability & Scale

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior SRE: Observability, Automation & Hybrid Cloud
Senior SRE: Observability, Automation & Hybrid Cloud

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 140,000 - 165,000
Hybrid work option
On-call rotation
Work from home eligibility
Senior SRE: Cloud Reliability, Observability Lead
Senior SRE: Cloud Reliability, Observability Lead

CentralReach • Fort Lauderdale (FL)

Hybrid
USD 160,000 - 180,000
Hybrid work model
Health benefits
PTO & 401(k) matching
+1
Senior Cloud SRE: Observability, Automation & Resilience
Senior Cloud SRE: Observability, Automation & Resilience

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior SRE: Hybrid Cloud Reliability & Observability Lead
Senior SRE: Hybrid Cloud Reliability & Observability Lead

Colorado PERA • Denver (CO)

Hybrid
USD 140,000 - 190,000
Hybrid work option