Site Reliability Engineer

100 CRC Insurance Group, LLC

Charlotte (NC)

On-site

USD 130,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with match
Generous PTO
Restricted stock units potential

Job summary

CRC Group seeks a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our observability strategy across cloud and on‑prem systems. This full‑time leadership role owns monitoring design, drives platform decisions, and guides engineering teams toward modern SRE practices.

You will define end‑to‑end monitoring pipelines, own Dynatrace and ServiceNow integration, and champion proactive reliability with SLIs/SLOs, alerting standards, and automated remediation.

Qualifications

  • 7+ years in Site Reliability Engineering, monitoring, or production engineering.
  • Proven experience in a technical leadership role.
  • Deep hands-on experience with Dynatrace, Azure, and ServiceNow ITSM/ITOM.
  • Ability to design and lead enterprise monitoring/SRE architectures and drive platform decisions.

Responsibilities

  • Define and own the enterprise monitoring and SRE observability strategy.
  • Serve as SME for Dynatrace, ServiceNow integration, and alerting architecture.
  • Architect end-to-end monitoring and SRE pipelines (Dynatrace → ServiceNow).
  • Lead adoption of SRE principles: SLIs, SLOs, and error budgets.
  • Drive automation of remediation and self-healing workflows.
  • Coordinate across engineering, cloud, and ITSM teams to standardize monitoring.

Skills

Dynatrace expertise
Azure IaaS/PaaS
ServiceNow ITSM/ITOM
SRE architecture design
Leadership
Automation frameworks
CI/CD tooling

Tools

Terraform
GitHub Actions
Runbooks
ARM/Bicep

Job description

Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)

We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This full‑time leadership role is responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern Site Reliability Engineering (SRE) practices.

Job Details
  • Type: Regular
  • Language Fluency: English (Required)
  • Work Shift: 1st Shift (United States of America)
Key Responsibilities
  • Strategic Leadership & Decision-Making
    • Define and own the enterprise monitoring and SRE observability strategy.
    • Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture.
    • Evaluate and recommend tooling, integration patterns, and platform direction.
    • Drive decisions on alerting philosophy, noise reduction, and signal quality improvement.
  • Platform Ownership & Architecture
    • Architect and standardize end‑to‑end monitoring and SRE pipelines: Dynatrace → ServiceNow incident lifecycle.
    • Implement alert correlation, deduplication, and prioritization.
    • Integrate with paging systems (PagerDuty, SMS, voice, Teams).
    • Establish best practices for event ingestion and enrichment, incident routing and automated assignment, CMDB and service mapping.
  • SRE Leadership
    • Lead adoption of SRE principles: SLIs, SLOs, and error budgets.
    • Promote reliability engineering practices across services.
    • Encourage proactive monitoring and resilience design.
    • Champion shift from reactive operations to proactive reliability engineering.
    • Influence application and platform teams to build observable, resilient systems by design.
  • Automation & Self‑Healing Enablement
    • Drive development of automated remediation and self‑healing capabilities.
    • Leverage Dynatrace workflows, Azure services, and automation frameworks to reduce manual incident handling, eliminate repeatable operational tasks, and minimize unnecessary paging.
  • ServiceNow & Observability Integration Leadership
    • Own integration between Dynatrace and ServiceNow ITSM/ITOM, including incident, event management, and CMDB alignment.
    • Define standards for automated incident creation and resolution, priority assignment and routing logic, and monitoring‑to‑ITSM data synchronization.
  • Team Leadership & Cross‑Functional Influence
    • Provide technical leadership and mentorship across SRE, platform, and application teams.
    • Act as a central point of coordination between engineering, cloud, and ITSM teams.
    • Lead workshops and working sessions to drive monitoring standardization, align teams on reliability practices, and influence upstream architectural decisions.
  • Operational Excellence
    • Establish KPIs and drive improvement in incident response and resolution times, alert quality and paging effectiveness, and monitoring coverage across critical services.
    • Provide leadership with clear visibility into service health and reliability trends.
Required Qualifications
  • 7+ years in Site Reliability Engineering, monitoring, or production engineering.
  • Proven experience in a technical leadership or lead engineer role.
  • Deep hands‑on experience with Dynatrace (or equivalent observability platforms), Microsoft Azure (IaaS, PaaS, networking, identity), and ServiceNow ITSM/ITOM.
  • Demonstrated ability to design and lead enterprise monitoring/SRE architectures, drive platform and tooling decisions, and integrate observability, ITSM, and paging solutions.
Preferred Qualifications
  • Experience leading SRE or observability transformation initiatives.
  • Strong expertise with Dynatrace–ServiceNow integrations.
  • Experience modernizing or consolidating paging/on‑call tooling.
  • Familiarity with Azure‑based SRE tooling or AI‑assisted operations.
  • Comfort with automation frameworks (GitHub Actions, Runbooks, etc.) and Infrastructure as Code (Terraform, ARM, Bicep).
  • Track record of achieving success metrics such as reduction in alert noise, improved incident routing accuracy, and increased adoption of self‑healing workflows.
  • Alignment between monitoring, CMDB, and service ownership across enterprise‑wide adoption of SRE and monitoring standards.
Benefits
  • Medical, dental, vision, life, disability, and AD&D insurance.
  • Tax‑advantaged savings accounts and a 401(k) plan with company match.
  • Generous paid time off, including company holidays, vacation, sick days, and new parent leave.
  • Potential eligibility for restricted stock units and/or a deferred compensation plan.
  • Drug‑free workplace.
Equal Opportunity Employer Statement

CRC Group is an Equal Opportunity Employer that does not discriminate based on race, gender, color, religion, citizenship, national origin, age, sexual orientation, gender identity, disability, veteran status, or any other classification protected by law. CRC Group supports a diverse workforce and is a Drug Free Workplace. EEO is the law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

CRC Group • Town of Charlotte (NY)

On-site
USD 140,000 - 190,000
401(k) with company match
Paid time off
Job health and wellness benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CRC Group • Charlotte (NC)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

CRC Group • Charlotte (NC)

On-site
USD 120,000 - 190,000
Medical, dental, vision insurance
401(k) plan with company match
Paid time off
+1
Lead Observability & SRE Architect
Lead Observability & SRE Architect

CRC Group • Charlotte (NC)

On-site
USD 130,000 - 180,000
Lead SRE & Observability Engineer (Azure, Dynatrace)
Lead SRE & Observability Engineer (Azure, Dynatrace)

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Lead SRE: Observability, Dynatrace & Automation
Lead SRE: Observability, Dynatrace & Automation

CRC Group • Charlotte (NC)

On-site
USD 120,000 - 190,000
Medical, dental, vision insurance
401(k) plan with company match
Paid time off
+1
Lead SRE & Observability Engineer - Dynatrace/ServiceNow
Lead SRE & Observability Engineer - Dynatrace/ServiceNow

CRC Group • Town of Charlotte (NY)

On-site
USD 140,000 - 190,000
401(k) with company match
Paid time off
Job health and wellness benefits
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert

RaceTrac Petroleum, Inc. • United States

Hybrid
USD 110,000 - 150,000
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert

RaceTrac • Atlanta (GA)

On-site
Senior SRE: Dynatrace & Azure Observability Expert
Senior SRE: Dynatrace & Azure Observability Expert

RaceTrac • Atlanta (GA)

On-site