SRE Specialist

TechDigital Group

Dallas (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A technology services firm in Dallas is seeking a skilled professional to manage Site Reliability Engineering (SRE) aspects. Candidates must have significant experience in Azure cloud operations and distributed systems, along with expertise in observability tools such as Dynatrace and Grafana. The role requires strong scripting and automation skills, as well as leadership and client-facing capabilities. This opportunity promises to challenge and grow your SRE skills in a fast-paced environment.

Qualifications

  • Strong experience in Azure cloud operations and large scale distributed systems.
  • Hands-on expertise with observability tools such as Dynatrace, Newrelic/Datadog, Prometheus, and Grafana.
  • Solid understanding of ITSM processes including incident, problem, and change management.
  • Strong scripting skills in Python and Shell for troubleshooting complex production issues.
  • Proficiency in automation tooling and CI/CD pipelines.

Responsibilities

  • Guide teams on SRE aspects and aspects of service reliability.
  • Implement and maintain disaster recovery and backup strategies.
  • Ensure strong leadership, communication, and client-facing capabilities.

Skills

Prometheus
Dynatrace
Grafana
Azure cloud operations
Python
Shell
Terraform
Ansible
GitOps

Job description

Skills
  • Prometheus
  • Dynatrace
  • Grafana
  • Azure cloud operations
Job Description
  • Must have good understanding of SRE aspects and guide teams
  • Strong experience in Azure cloud operations and large scale distributed systems.
  • Hands on expertise with observability tools: Dynatrace, Newrelic/Datadog, Prometheus, and Grafana.
  • Solid understanding of ITSM processes (incident, problem, change).
  • Strong scripting and troubleshooting skills (Python, Shell) for complex production issues.
  • Implement and maintain disaster recovery, failover mechanism and backup strategies
  • Proficiency in automation tooling: Terraform, Ansible, GitOps, and CI/CD pipelines.
  • Strong leadership, communication, cross functional collaboration, and client facing skills
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineer
SRE Engineer

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Senior SRE - Azure
Senior SRE - Azure

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 100,000 - 140,000
Azure SRE Leader: Observability, Automation & Resilience
Azure SRE Leader: Observability, Automation & Resilience

TechDigital Group • Dallas (TX)

On-site
USD 100,000 - 130,000
SRE Lead
SRE Lead

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Solutions Architect(SRE)
Solutions Architect(SRE)

Business Needs Inc. • Fort Mill (SC)

On-site
USD 100,000 - 130,000
Senior SRE (Site Reliability Engineer)
Senior SRE (Site Reliability Engineer)

Vytwo • Dallas (TX)

Hybrid
USD 130,000 - 160,000
Flexible work from home options
SRE Manager
SRE Manager

Mphasis • New Jersey

On-site
USD 170,000 - 180,000