Senior Site Reliability Engineer

CRC Group

Charlotte (NC)

On-site

USD 130,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CRC Group is seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This full-time leadership role owns monitoring design, drives platform decisions, and guides engineering teams toward modern SRE practices.

You will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and paging systems, enabling automated,

Qualifications

  • Proven leadership in SRE or platform reliability roles.
  • Deep hands-on experience with Dynatrace and ServiceNow ITSM/ITOM.
  • Ability to design enterprise monitoring/SRE architectures.
  • Experience integrating observability, ITSM, and paging solutions.
  • Familiar with cloud tooling and automation frameworks.

Responsibilities

  • Define and own the enterprise monitoring and SRE observability strategy.
  • Lead Dynatrace–ServiceNow integration and alerting architecture.
  • Drive tooling decisions and platform direction across teams.
  • Architect end-to-end monitoring pipelines and incident workflows.
  • Provide technical leadership and mentor SRE, platform, and app teams.
  • Champion automated remediation and self-healing capabilities.
  • Align monitoring with CMDB and service mapping, improve paging.

Skills

Dynatrace
ServiceNow ITSM/ITOM
SRE leadership
Monitoring architecture
Alerting & paging
Cloud platforms
Automation
Azure

Tools

Terraform
ARM/Bicep
PagerDuty

Job description

We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full‑time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.

This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self‑healing operations.

Location: This role is based in Charlotte, NC.

Key Responsibilities
Strategic Leadership & Decision‑Making
  • Define and own the enterprise monitoring and SRE observability strategy
  • Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture
  • Evaluate and recommend tooling, integration patterns, and platform direction
  • Drive decisions on alerting philosophy, noise reduction, and signal quality improvement
  • Architect and standardize end‑to‑end monitoring and SRE pipelines:
  • Dynatrace → ServiceNow incident lifecycle
  • Alert correlation, deduplication, and prioritization
  • Integration with paging systems (PagerDuty, SMS, voice, Teams)
  • Establish best practices for:
  • Event ingestion and enrichment
  • Incident routing and automated assignment
  • Integration with CMDB and service mapping
Site Reliability Engineering (SRE) Leadership
  • Lead adoption of SRE principles, including:
  • SLIs, SLOs, and error budgets
  • Reliability engineering practices across services
  • Proactive monitoring and resilience design
  • Champion a shift from reactive operations to proactive reliability engineering
  • Influence application and platform teams to build observable, resilient systems by design
Automation & Self‑Healing Enablement
  • Drive development of automated remediation and self‑healing capabilities
  • Leverage Dynatrace workflows, Azure services, and automation frameworks to:
  • Reduce manual incident handling
  • Minimize unnecessary paging
ServiceNow & Observability Integration Leadership
  • Own integration between Dynatrace and ServiceNow ITSM/ITOM, including:
  • Incident, Event Management, and CMDB alignment
  • Service mapping and dependency visibility
  • Define standards for:
  • Automated incident creation and resolution
  • Priority assignment and routing logic
Team Leadership & Cross‑Functional Influence
  • Provide technical leadership and mentorship across SRE, platform, and application teams
  • Act as a central point of coordination between engineering, cloud, and ITSM teams
  • Lead workshops and working sessions to:
  • Align teams on reliability practices
Operational Excellence
  • Establish KPIs and drive improvement in:
  • Incident response and resolution times
  • Alert quality and paging effectiveness
  • Monitoring coverage across critical services
  • Provide leadership with clear visibility into service health and reliability trends
Required Qualifications
  • 7+ years in Site Reliability Engineering, monitoring, or production engineering
  • Proven experience in a technical leadership or lead engineer role
  • Deep hands‑on experience with:
  • Dynatrace (or equivalent observability platforms)
  • ServiceNow ITSM / ITOM (incident, event management, CMDB)
  • Demonstrated ability to:
  • Design and lead enterprise monitoring/SRE architectures
  • Drive platform and tooling decisions
  • Integrate observability, ITSM, and paging solutions
Preferred Qualifications
  • Experience leading SRE or observability transformation initiatives
  • Strong expertise with Dynatrace–ServiceNow integrations
  • Experience modernizing or consolidating paging/on‑call tooling
  • Familiarity with:
  • Azure‑based SRE tooling or AI‑assisted operations
  • Infrastructure as Code (Terraform, ARM, Bicep)
Success Metrics
  • Reduction in alert noise and unnecessary paging
  • Improved incident routing accuracy and MTTR
  • Increased adoption of self‑healing and automated workflows
  • Strong alignment between monitoring, CMDB, and service ownership
  • Enterprise‑wide adoption of SRE and monitoring standards
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Lead Observability & SRE Architect
Lead Observability & SRE Architect

CRC Group • Charlotte (NC)

On-site
USD 130,000 - 180,000
Senior SRE (Site Reliability Engineer)
Senior SRE (Site Reliability Engineer)

Vytwo • Dallas (TX)

Hybrid
USD 130,000 - 160,000
Flexible work from home options
Lead SRE: Observability, Dynatrace & Automation
Lead SRE: Observability, Dynatrace & Automation

CRC Group • Charlotte (NC)

On-site
USD 120,000 - 190,000
Medical, dental, vision insurance
401(k) plan with company match
Paid time off
+1
Lead SRE & Observability Engineer (Azure, Dynatrace)
Lead SRE & Observability Engineer (Azure, Dynatrace)

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Site Reliability Engineer
Site Reliability Engineer

CRC Group • Charlotte (NC)

On-site
USD 120,000 - 190,000
Medical, dental, vision insurance
401(k) plan with company match
Paid time off
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer - Observability & Dynatrace
Site Reliability Engineer - Observability & Dynatrace

RaceTrac • Atlanta (GA)

Hybrid
USD 130,000 - 175,000