SRE-Azure - Global Industrial

Genuine Parts Company

Alabama

On-site

USD 110,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
401(k)
Tuition reimbursement
Vacation
Sick
Holiday pay

Job summary

Genuine Parts Company is seeking an Azure Site Reliability Engineer to improve reliability, availability, and performance of enterprise applications hosted in Microsoft Azure. This role blends software engineering, cloud infrastructure, automation, and DevOps to reduce toil while building resilient, cloud-native solutions.

You will monitor and optimize Azure-based platforms, define SLI/SLO/SLA, and partner with development teams to improve deployment automation, incident response, and platform

Qualifications

  • Bachelor's degree in technology or related field.
  • 5-7 years of experience in a technology or software engineering role.
  • Experience supporting large-scale, highly available, distributed applications and cloud platforms.
  • Strong knowledge of Site Reliability Engineering principles, incident management, and operational excellence.
  • Proficient in Azure cloud services, AKS, IaC, and observability tools.

Responsibilities

  • Monitor and optimize Azure-hosted platforms for performance and reliability.
  • Define SLIs, SLOs, SLAs, and error budgets.
  • Lead incident response and post-incident RCAs.
  • Automate deployments and reduce toil with IaC.
  • Design and support AKS, Azure Networking, App Services, and Storage.
  • Collaborate with cross-functional teams on reliability improvements.

Skills

SRE principles
Azure services
AKS
Terraform
ARM Templates
Azure DevOps
GitHub Actions
CI/CD
Monitoring
Windows Server
Linux
Hybrid cloud
Capacity planning
Disaster recovery
Business continuity
Analytical skills
Communication

Education

Bachelor's degree in technology or related field

Tools

Terraform
ARM Templates
AKS
Azure Monitor
Grafana
Datadog
Dynatrace
Application Insights
Log Analytics
Azure DevOps
GitHub Actions

Job description

The Azure Site Reliability Engineer (SRE) is responsible for improving the reliability, availability, scalability, performance, and operational excellence of enterprise applications and platforms hosted within Microsoft Azure. This role combines software engineering, cloud infrastructure, automation, and DevOps practices to build and support resilient, cloud-native solutions while reducing operational toil through automation.

The Azure SRE leverages Azure platform services, Kubernetes, Infrastructure as Code (IaC), and observability tools to ensure mission-critical systems remain highly available, secure, and performant. This role partners closely with application development, cloud engineering, architecture, and cybersecurity teams to drive continuous improvement, accelerate cloud adoption, and maintain service reliability through proactive monitoring, incident management, and operational engineering.

JOB DUTIES
  • Monitor, analyze, and optimize system performance, availability, and reliability across Azure-hosted platforms and applications.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.
  • Drive continuous service improvement through operational metrics, trend analysis, reliability engineering practices, and platform modernization efforts.
  • Partner with development teams to improve service reliability through testing, release validation, deployment automation, and production readiness reviews.
  • Design, build, and support Azure infrastructure and platform services, including Azure Kubernetes Service (AKS), Azure Networking, App Services, and Storage
  • Develop and maintain Infrastructure as Code (IaC) solutions utilizing Terraform and Azure-native deployment technologies.
  • Lead and support incident response, root cause analysis (RCA), post-incident reviews, and service restoration efforts.
  • Automate operational processes, platform provisioning, deployments, and remediation activities to reduce manual effort (TOIL) and improve reliability.
  • Identify, investigate, and mitigate performance, security, networking, and availability issues, including traffic anomalies and service disturbances.
  • Participate in on-call rotations and maintain operational documentation, runbooks, and knowledge articles as needed.
EDUCATION & EXPERIENCE
  • Typically requires a bachelor's degree and five (5) to seven (7) years of experience in a technology and/or software engineering role or an equivalent combination.
  • Experience supporting large-scale, highly available, distributed applications and cloud platforms.
KNOWLEDGE, SKILLS, ABILITIES
  • Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis, and Operational Excellence.
  • Hands-on experience with Microsoft Azure services, including infrastructure, networking, security, platform services, and cloud-native technologies.
  • Experience administering and supporting Azure Kubernetes Service (AKS), Kubernetes clusters, containers, and scalable distributed systems.
  • Proficiency with Infrastructure as Code (Terraform, ARM Templates) and Git-based deployment practices.
  • Experience with Azure DevOps, GitHub Actions, CI/CD pipelines, release automation, and DevOps methodologies.
  • Strong troubleshooting skills across cloud infrastructure, operating systems, databases, networking, and security domains.
  • Experience with monitoring and observability platforms, including Azure Monitor, Log Analytics, Application Insights, Grafana, Datadog, and Dynatrace.
  • Knowledge of microservices, APIs, distributed architectures, and cloud-native design patterns.
  • Working knowledge of Windows Server, Linux, networking, DNS, load balancing, firewalls, and hybrid cloud connectivity.
  • Experience with capacity planning, performance engineering, scalability testing, disaster recovery, and business continuity practices.
  • Strong analytical, problem-solving, communication, and collaboration skills.
BUDGET RESPONSIBILITY:

No

COMPANY INFORMATION:

Motion offers an excellent benefits package which includes options for healthcare coverage, 401(k), tuition reimbursement, vacation, sick, and holiday pay.

  • Healthcare coverage
  • 401(k)
  • Tuition reimbursement
  • Vacation
  • Sick
  • Holiday pay
DISCLAIMER:

This job description illustrates the general nature and level of work performed by employees within this job classification. It is not intended to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and skills required. Management retains the right to add or modify duties at any time.

GPC conducts its business without regard to sex, race, creed, color, religion, marital status, national origin, citizenship status, age, pregnancy, sexual orientation, gender identity or expression, genetic information, disability, military status, status as a veteran, or any other protected characteristic. GPC's policy is to recruit, hire, train, promote, assign, transfer and terminate employees based on their own ability, achievement, experience and conduct and other legitimate business reasons.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE-Azure - Global Industrial
SRE-Azure - Global Industrial

Motion Industries (MOT) • Alabama

On-site
USD 110,000 - 160,000
Healthcare coverage
401(k) plan
Tuition reimbursement
SRE-Azure - Global Industrial
SRE-Azure - Global Industrial

USA MOT Motion Industries, Inc. • Alton (IL)

On-site
USD 120,000 - 170,000
Healthcare coverage
401(k)
Tuition reimbursement
+3
SRE-Azure - Global Industrial
SRE-Azure - Global Industrial

Motion • Birmingham (AL), Northern (KY)

Hybrid
USD 110,000 - 170,000
Healthcare coverage
401(k)
Tuition reimbursement
+3
Site Reliability Engineer III
Site Reliability Engineer III

USA MOT Motion Industries, Inc. • Alton (IL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer II
Site Reliability Engineer II

Talentify • Birmingham (AL)

On-site
USD 90,000 - 120,000
M365/Azure-Sys Admin ll - Global Industrial
M365/Azure-Sys Admin ll - Global Industrial

Genuine Parts Company • Alabama

On-site
USD 85,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Site Reliability Engineer II
Site Reliability Engineer II

USA MOT Motion Industries, Inc. • Alton (IL)

On-site
USD 110,000 - 160,000
M365/Azure-Sys Admin ll - Global Industrial
M365/Azure-Sys Admin ll - Global Industrial

USA MOT Motion Industries, Inc. • United States

On-site
USD 90,000 - 130,000
Healthcare coverage
401(k) program
Tuition reimbursement
+1
Site Reliability Engineer - ED&A - Global Industrial
Site Reliability Engineer - ED&A - Global Industrial

USA MOT Motion Industries, Inc. • Alton (IL)

On-site
USD 110,000 - 140,000
Healthcare coverage
401(k)
Tuition reimbursement
+2