Site Reliability Engineer - Observability & Dynatrace

RaceTrac

Atlanta (GA)

Hybrid

USD 130,000 - 175,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RaceTrac in Atlanta, GA is seeking a highly experienced Site Reliability Engineer (SRE) with deep expertise in Dynatrace, observability engineering, and Azure cloud technologies. The role focuses on building and managing enterprise observability across critical digital platforms.

Ideal candidates will have hands-on Dynatrace (DQL) mastery, Azure Monitor, App Insights, Azure Functions, and APIM, plus proficiency reading .NET code to troubleshoot and implement observability in Azure environments

Qualifications

  • 10+ years IT experience with enterprise systems.
  • Expert in Dynatrace and DQL.
  • Azure cloud, App Insights, and Azure Functions experience.

Responsibilities

  • Design and optimize enterprise observability using Dynatrace.
  • Develop Dynatrace queries, dashboards, and alerts.
  • Monitor Azure services, APIs, and distributed systems.
  • Collaborate with engineering and cloud teams to improve reliability.

Skills

Dynatrace
DQL
Observability
Azure cloud
.NET code
Mobile observability
Telemetry & logging

Tools

Dynatrace
Azure Monitor
KQL
App Insights
APIM

Job description

Location: Hybrid (2-3 days in office) in Atlanta, GA (Cumberland Mall-Galleria)

We are seeking a highly experienced Site Reliability Engineer (SRE) with deep expertise in Dynatrace, observability engineering, and Azure cloud technologies. This role will be exclusively focused on building, enhancing, and managing enterprise observability, telemetry, monitoring, and proactive reliability engineering practices across critical digital platforms.

The ideal candidate must possess advanced hands-on expertise in Dynatrace, especially Dynatrace Query Language (DQL), along with strong knowledge of Azure Monitor, Azure KQL, Application Insights, Azure Functions, APIM, and distributed telemetry concepts. The candidate should have a strong understanding of .NET application architecture and the ability to read and analyze .NET code to support troubleshooting, root cause analysis, and observability implementation within Azure environments. Experience enabling observability for mobile platforms such as iOS and Android is also required.

This is a highly technical, hands-on role requiring a proactive engineering mindset, strong analytical capabilities, and the ability to collaborate across engineering, cloud, mobile, and business teams.

What You'll Do
Dynatrace & Observability Engineering
  • Serve as the primary Dynatrace SME across the organization.
  • Design, develop, and optimize enterprise observability solutions using Dynatrace.
  • Develop advanced Dynatrace DQL queries, dashboards, workflows, alerts, and analytics.
  • Implement intelligent monitoring strategies for applications, APIs, integrations, Azure services, mobile platforms, and distributed systems.
  • Continuously improve observability maturity through telemetry standardization, proactive monitoring, and automation.
  • Configure and tune alerting mechanisms to improve signal-to-noise ratio and reduce alert fatigue.
  • Leverage Dynatrace Davis AI, anomaly detection, and AI-driven root cause analysis capabilities.
  • Enable and enhance observability for mobile applications across iOS and Android platforms.
Azure Monitoring & Cloud Operations
  • Build and maintain monitoring solutions using:
  • Application Insights
  • Monitor and troubleshoot Azure Function Apps, App Services, APIs, integrations, and backend services.
  • Analyze telemetry, traces, logs, metrics, and distributed transactions to identify root causes and performance bottlenecks.
  • Troubleshoot cloud-native applications and Azure infrastructure issues.
  • Develop proactive monitoring for cloud services, integrations, APIs, and backend processing systems.
API & Integration Monitoring
  • Monitor and troubleshoot Azure API Management (APIM), API Gateways, API endpoints, and integrations.
  • Understand end-to-end API transaction flows and dependency mapping.
  • Build observability solutions for APIs, middleware platforms, and integration services.
  • Diagnose latency issues, transaction failures, authentication issues, and backend service degradation.
  • Enable telemetry, monitoring, tracing, and performance analysis for iOS and Android applications.
  • Analyze mobile-to-backend transaction flows and end-user experience metrics.
  • Troubleshoot mobile application latency, crash analytics, API failures, and connectivity issues.
  • Correlate mobile telemetry with backend application and infrastructure monitoring data.
Application Engineering & Troubleshooting
  • Utilize prior .NET development experience to troubleshoot application behavior, performance, and deployment issues.
  • Read and understand .NET application code to support root cause analysis and observability implementation.
  • Work closely with development teams to understand application logic, API flows, dependencies, and exception handling.
  • Support Azure Function deployments, configuration management, scaling, and runtime troubleshooting.
  • Collaborate with development teams during architecture reviews and production releases.
  • Ensure observability and monitoring readiness before deployments go live.
Site Reliability Engineering (SRE)
  • Perform deep technical analysis across systems by correlating logs, metrics, traces, and application telemetry.
  • Conduct root cause analysis (RCA) for recurring incidents and systemic issues.
  • Partner with engineering and operations teams to implement preventive improvements and automation.
  • Develop KPI-driven reliability improvements focused on system stability, performance, and operational excellence.
  • Proactively identify risks, bottlenecks, failure patterns, and reliability concerns before business impact occurs.
  • Automate operational workflows and monitoring processes wherever possible.
  • Improve operational efficiency using AI-driven insights and automation capabilities.
  • Build reusable monitoring frameworks, dashboards, and telemetry standards.
  • Drive observability best practices across engineering teams.
What We're Looking For
Mandatory Technical Skills
  • 10+ years of overall IT experience.
  • Expert-level hands-on experience with Dynatrace.
  • Advanced expertise in Dynatrace Query Language (DQL).
  • Deep understanding of telemetry, observability, distributed tracing, metrics, and logging concepts.
  • Strong Azure cloud experience with emphasis on:
  • Application Insights
  • Strong understanding of API architectures, API Gateways, and backend integrations.
  • Strong ability to read, analyze, and understand .NET application code.
  • Experience troubleshooting and deploying Azure Functions and cloud-native applications.
  • Experience enabling observability and telemetry for mobile applications on iOS and Android.
  • Understanding of mobile telemetry, crash analytics, API monitoring, and end-user experience monitoring.
  • Strong understanding of distributed systems and enterprise application architectures.
Preferred Skills
  • Experience with OpenTelemetry implementation and instrumentation.
  • Experience with CI/CD pipelines and DevOps practices.
  • Knowledge of AI-driven observability and AIOps concepts.
  • Familiarity with ServiceNow and incident management workflows.
  • Experience with Databricks, SQL platforms, and integration technologies.
Core Competencies
  • Strong analytical and troubleshooting skills.
  • Excellent communication and stakeholder management abilities.
  • Ability to work independently and drive proactively.
  • Strong collaboration skills across engineering, cloud, SRE, mobile, and business teams.
  • Ability to quickly adapt to new technologies and evolving environments.
Success Criteria
  • Reduction in recurring incidents through proactive monitoring and RCA.
  • Improved observability coverage across enterprise systems, APIs, and mobile applications.
  • Faster incident detection and resolution.
  • Reduction in monitoring noise and false positives.
  • Increased automation and operational efficiency.
  • Improved reliability and performance of critical systems and APIs.
  • Strong partnership with engineering teams to ensure production readiness and operational excellence.
#Atlanta
Fueled by Growth, Driven by You

At RaceTrac, our people make the difference. Whether you’re working in a store, at our corporate office, or on the road, you’ll be part of a team that brings energy, innovation, and a passion for serving others every day. We support each other, celebrate wins big and small, and create opportunities for growth at every level. With four operating divisions RaceTrac, RaceWay, Energy Dispatch, and Gulf - there’s always a new challenge to take on and a new path to pursue. Join us and discover how far your career can go.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert

RaceTrac Petroleum, Inc. • United States

Hybrid
USD 110,000 - 150,000
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert
Senior Site Reliability Engineer (SRE) – Dynatrace & Azure Observability Expert

3M HEALTHCARE • Atlanta (GA)

Hybrid
USD 140,000 - 180,000
Senior SRE: Dynatrace & Azure Observability Expert
Senior SRE: Dynatrace & Azure Observability Expert

3M HEALTHCARE • Atlanta (GA)

Hybrid
USD 140,000 - 180,000
Senior SRE: Dynatrace Obs & Azure Telemetry Architect
Senior SRE: Dynatrace Obs & Azure Telemetry Architect

RaceTrac • Atlanta (GA)

Hybrid
USD 130,000 - 175,000
Senior SRE: Dynatrace & Azure Observability Architect
Senior SRE: Dynatrace & Azure Observability Architect

RaceTrac Petroleum, Inc. • United States

Hybrid
USD 110,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Senior SRE (Site Reliability Engineer)
Senior SRE (Site Reliability Engineer)

Vytwo • Dallas (TX)

Hybrid
USD 130,000 - 160,000
Flexible work from home options
Solutions Engineer - FED CIV - (REMOTE DMV)
Solutions Engineer - FED CIV - (REMOTE DMV)

Dynatrace LLC • Washington, Northern (KY)

Hybrid
USD 100,000 - 125,000
Epic Site Reliability Engineer II
Epic Site Reliability Engineer II

Quest Diagnostics Incorporated • Schaumburg (IL)

Hybrid
USD 85,000 - 110,000
Medical benefits
401(k) match
Employee stock purchase plan
+1
Epic Site Reliability Engineer II
Epic Site Reliability Engineer II

Quest Diagnostics • Secaucus (NJ)

Hybrid
USD 85,000 - 110,000
Medical benefits
401(k) match
Education assistance