Site Reliability Engineer III

Genuine Parts Company

Alabama

On-site

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare coverage
401(k)
Tuition reimbursement
Vacation, sick, and holiday pay

Job summary

Genuine Parts Company is seeking a Site Reliability Engineer III in Alabama. This role focuses on improving system reliability and resilience while building automation to prevent incidents. Responsibilities include system performance metrics analysis, working with development teams on service improvements, and overseeing cloud-based transformations. Ideal candidates will have a bachelor's degree and at least five years of relevant experience, along with expertise in Kubernetes and cloud services. The position offers competitive benefits including healthcare and a 401(k).

Qualifications

  • Typically requires a bachelor's degree and five or more years of related experience.
  • Strong software engineering and operational support skills.

Responsibilities

  • Gathers and analyzes metrics from monitoring platforms for performance tuning.
  • Partners with development teams to improve services through testing.
  • Participates in system design, platform management and capacity planning.
  • Balances feature development speed and reliability.
  • Works closely with the incident response team.

Skills

Understanding of Kubernetes, containers, clusters, and elastic scalability
Expertise in SRE principles
Cloud Services experience with Google Cloud Platform (GCP)
Experience with API, service-based or microservice-based architecture
Proficiency in infrastructure, network, database, operating systems, or security troubleshooting
Experience with Azure DevOps (ADO), Dynatrace, Prometheus, Terraform and Grafana

Education

Bachelor's degree in a related field

Tools

Dynatrace
Prometheus
Terraform
Grafana

Job description

Site Reliability Engineer III

Under limited supervision, the Site Reliability Engineer III is responsible for improving system reliability and resilience. This role focuses on building automation to reduce manual effort and prevent service-impacting incidents. The SRE combines software and systems engineering to build and support large-scale, distributed, fault-tolerant systems. This role ensures that critical platforms are available, reliable, and able to support a fast rate of improvement. This role relies on monitoring platforms and is continually taking a holistic view of system health and performance. The SRE will enhance and support cloud-based transformations and is focused on pushing capabilities forward, staying ahead of customer needs, and innovating for continuous improvement. The SRE provides operational support and engineering for multiple large-scale distributed software applications.

You must be eligible to work in the US without Visa Sponsorship.

Job Duties
  • Gathers and analyzes metrics from monitoring platforms to assist in performance tuning and fault tolerance.
  • Partners with development teams to improve services through testing and release procedures.
  • Participates in system design, platform management and capacity planning.
  • Balances feature development speed and reliability with service-level objectives.
  • Works closely with the incident response team and restoring service to normal operation.
  • Understands debugging and applying troubleshooting skills.
  • Investigates, blocks and rate-limits unwanted traffic.
  • Utilizes monitoring systems and dashboards for proactive changes and alerting.
  • Establishes continuous process improvement cycles where the process, performance, and supporting technologies are reviewed and enhanced where applicable.
  • Performs other duties as assigned.
Education & Experience

Typically requires a bachelor's degree and five (5) or more years of related experience or an equivalent combination.

Knowledge, Skills, Abilities
  • Understanding of Kubernetes, containers, clusters, and elastic scalability.
  • Expertise in SRE principles.
  • Mindset of continually finding ways to drive scalability, stability, and performance.
  • Cloud Services experience with Google Cloud Platform (GCP).
  • Experience with API, service-based or microservice-based architecture.
  • Proficiency in infrastructure, network, database, operating systems, or security troubleshooting and remediation.
  • Architecture-level knowledge of Windows and Linux and Infrastructure systems.
  • Experience with production deployment, monitoring, and operational support for enterprise-class applications (Dynatrace a plus).
  • Experience working with Continuous Integration/ Continuous Deployment tools.
  • Experience in performance diagnostics, capacity planning, performance architecture design, performance tuning, and performance monitoring.
  • A strong mix of software engineering and operational support skills.
  • Knowledge of web technologies – HTTP, proxy, java, etc.
  • Experience with Azure DevOps (ADO), Dynatrace, Prometheus, Terraform and Grafana.
Supervisory Responsibility

No Supervisory Responsibility

Budget Responsibility

No

Company Information

Motion offers an excellent benefits package which includes options for healthcare coverage, 401(k), tuition reimbursement, vacation, sick, and holiday pay.

Disclaimer

This job description illustrates the general nature and level of work performed by employees within this job classification. It is not intended to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and skills required. Management retains the right to add or modify duties at any time.

GPC conducts its business without regard to sex, race, creed, color, religion, marital status, national origin, citizenship status, age, pregnancy, sexual orientation, gender identity or expression, genetic information, disability, military status, status as a veteran, or any other protected characteristic. GPC's policy is to recruit, hire, train, promote, assign, transfer and terminate employees based on their own ability, achievement, experience and conduct and other legitimate business reasons.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Genuine Parts Company • Alabama

On-site
USD 80,000 - 110,000
Healthcare coverage
401(k)
Tuition reimbursement
+3
Manager of Site Reliability Engineering (SRE)
Manager of Site Reliability Engineering (SRE)

Genuine Parts Company • Alabama

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer II
Site Reliability Engineer II

Motion • Birmingham (AL)

On-site
USD 80,000 - 110,000
Healthcare coverage
401(k)
Tuition reimbursement
+3
Principal III, SRE
Principal III, SRE

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Experis • Charlotte (NC)

Hybrid
USD 96,000 - 103,000
Cloud Infrastructure Site Reliability Engineer
Cloud Infrastructure Site Reliability Engineer

Robotics Prcocess Automation, LLC • Berkeley Heights (NJ)

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NAB Leadership Foundation • San Antonio (TX)

On-site
USD 90,000 - 120,000
Employer sponsored medical, dental, and vision coverage
Paid vacation and sick time
401K plan
+1
Manager of Site Reliability Engineering (SRE)
Manager of Site Reliability Engineering (SRE)

Motion • Birmingham (AL)

Hybrid
USD 130,000 - 160,000