Site Reliability Engineer

CRC Group

Charlotte (NC)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Life insurance
Disability insurance
AD&D insurance
401(k) with company match
Paid time off
New parent leave
Restricted stock units
Deferred compensation

Job summary

CRC Group seeks a Lead Site Reliability & Environment Monitoring Engineer to define and drive our observability strategy across cloud and application platforms. You will own monitoring design, guide engineering teams, and advance proactive, self-healing operations.

This leadership role requires deep hands-on experience with Dynatrace, ServiceNow ITSM/ITOM, and Azure, with a focus on reducing noise and improving incident routing across our enterprise estate.

Qualifications

  • 7+ years in Site Reliability Engineering, monitoring, or production engineering.
  • Proven experience in a technical leadership or lead engineer role.
  • Deep hands-on experience with Dynatrace, Azure, and ServiceNow ITSM/ITOM.

Responsibilities

  • Define and own the enterprise monitoring and SRE observability strategy.
  • Serve as SME for Dynatrace, ServiceNow integration, and alerting architecture.
  • Evaluate tooling, integration patterns, and platform direction.
  • Drive decisions on alerting philosophy and signal quality improvements.
  • Architect end-to-end monitoring pipelines (Dynatrace→ServiceNow lifecycle).
  • Lead adoption of SRE principles including SLIs, SLOs, and error budgets.
  • Develop automated remediation and self-healing capabilities.

Skills

Dynatrace
Microsoft Azure
ServiceNow ITSM/ITOM
SRE Leadership

Job description

If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).

Regular or Temporary:

Regular

Language Fluency: English (Required)

Work Shift:

1st Shift (United States of America)

Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)

We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.

This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.

Location: This role is hybrid based in Charlotte, NC.

Strategic Leadership & Decision-Making
  • Define and own the enterprise monitoring and SRE observability strategy
  • Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture
  • Evaluate and recommend tooling, integration patterns, and platform direction
  • Drive decisions on alerting philosophy, noise reduction, and signal quality improvement
Platform Ownership & Architecture
  • Architect and standardize end-to-end monitoring and SRE pipelines:
    • Dynatrace → ServiceNow incident lifecycle
    • Alert correlation, deduplication, and prioritization
    • Integration with paging systems (PagerDuty, SMS, voice, Teams)
  • Establish best practices for:
    • Event ingestion and enrichment
    • Incident routing and automated assignment
    • Integration with CMDB and service mapping
Site Reliability Engineering (SRE) Leadership
  • Lead adoption of SRE principles, including:
    • SLIs, SLOs, and error budgets
    • Reliability engineering practices across services
    • Proactive monitoring and resilience design
  • Champion a shift from reactive operations to proactive reliability engineering
  • Influence application and platform teams to build observable, resilient systems by design
Automation & Self-Healing Enablement
  • Drive development of automated remediation and self-healing capabilities
  • Leverage Dynatrace workflows, Azure services, and automation frameworks to:
    • Reduce manual incident handling
    • Eliminate repeatable operational tasks
    • Minimize unnecessary paging
ServiceNow & Observability Integration Leadership
  • Own integration between Dynatrace and ServiceNow ITSM/ITOM, including:
    • Incident, Event Management, and CMDB alignment
    • Service mapping and dependency visibility
    • Governance for application/service tagging
  • Define standards for:
    • Automated incident creation and resolution
    • Priority assignment and routing logic
    • Monitoring-to-ITSM data synchronization
Team Leadership & Cross-Functional Influence
  • Provide technical leadership and mentorship across SRE, platform, and application teams
  • Act as a central point of coordination between engineering, cloud, and ITSM teams
  • Lead workshops and working sessions to:
    • Drive monitoring standardization
    • Align teams on reliability practices
    • Influence upstream architectural decisions
Operational Excellence
  • Establish KPIs and drive improvement in:
    • Incident response and resolution times
    • Alert quality and paging effectiveness
    • Monitoring coverage across critical services
  • Provide leadership with clear visibility into service health and reliability trends
Required Qualifications
  • 7+ years in Site Reliability Engineering, monitoring, or production engineering
  • Proven experience in a technical leadership or lead engineer role
  • Deep hands-on experience with:
    • Dynatrace (or equivalent observability platforms)
    • Microsoft Azure (IaaS, PaaS, networking, identity)
    • ServiceNow ITSM / ITOM (incident, event management, CMDB)
Preferred Qualifications
  • Experience leading SRE or observability transformation initiatives
  • Strong expertise with Dynatrace–ServiceNow integrations
  • Experience modernizing or consolidating paging/on-call tooling
  • Familiarity with:
    • Azure-based SRE tooling or AI-assisted operations
    • Automation frameworks (GitHub Actions, Runbooks, etc.)
    • Infrastructure as Code (Terraform, ARM, Bicep)
Success Metrics
  • Reduction in alert noise and unnecessary paging
  • Improved incident routing accuracy and MTTR
  • Increased adoption of self-healing and automated workflows
  • Strong alignment between monitoring, CMDB, and service ownership
  • Enterprise-wide adoption of SRE and monitoring standards

General Description of Available Benefits for Eligible Employees of CRC Group: At CRC Group, we're committed to supporting every aspect of teammates' well-being – physical, emotional, financial, social, and professional. Our best-in-class benefits program is designed to care for the whole you, offering a wide range of coverage and support. Eligible full‑time teammates enjoy access to medical, dental, vision, life, disability, and AD&D insurance; tax‑advantaged savings accounts; and a 401(k) plan with company match. CRC Group also offers generous paid time off programs, including company holidays, vacation and sick days, new parent leave, and more. Eligible positions may also qualify for restricted stock unitsand/or a deferred compensation plan.

CRC Group supports a diverse workforce and is an Equal Opportunity Employer that does not discriminate against individuals on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status or other classification protected by law. CRC Group is a Drug Free Workplace.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CRC Group • Charlotte (NC)

On-site
USD 130,000 - 180,000
Lead Observability & SRE Architect
Lead Observability & SRE Architect

CRC Group • Charlotte (NC)

On-site
USD 130,000 - 180,000
Lead SRE & Observability Engineer (Azure, Dynatrace)
Lead SRE & Observability Engineer (Azure, Dynatrace)

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Technical Operations Lead
Technical Operations Lead

First Citizens • Raleigh (NC)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

Ripple • New York (NY)

On-site
USD 130,000 - 180,000
Lead SRE – Azure, Dynatrace & ServiceNow
Lead SRE – Azure, Dynatrace & ServiceNow

CRC Group • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+8
Technical Operations Lead
Technical Operations Lead

First Citizens Bank • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Benefits program
Technical Operations Lead
Technical Operations Lead

First Citizens Bank • Dallas (TX)

On-site
USD 140,000 - 180,000
Senior SRE (Site Reliability Engineer)
Senior SRE (Site Reliability Engineer)

Vytwo • Dallas (TX)

On-site
USD 130,000 - 160,000
Flexible work from home options