A technology solutions company located in Atlanta, GA is seeking a Site Reliability Engineering (SRE) Architect. In this pivotal role, you will lead the design and implementation of resilient infrastructure and practices to ensure operational efficiency and reliability across services. The ideal candidate will have extensive experience in cloud computing, specifically AWS, along with a strong background in SRE principles and automation tools. This is a long-term position with a focus on developing best practices and mentoring others in the organization.
Qualifications
Proven experience in an architectural role focusing on reliability and performance.
Understanding of SRE principles like SLIs/SLOs and automation.
Strong experience with AWS and cloud architecture.
Responsibilities
Architect and design scalable infrastructure patterns on AWS.
Define SRE best practices and standards for service design.
Lead postmortems for significant incidents to improve system resilience.
Skills
Cloud computing expertise (AWS)
Containerization (Kubernetes, Docker)
Observability solutions design
Programming (Python, Go, Bash)
Strong analytical and problem-solving skills
Leadership and communication skills
Tools
Dynatrace
Prometheus
Grafana
ELK/EFK Stack
Jaeger
OpenTelemetry
Job description
A technology solutions company located in Atlanta, GA is seeking a Site Reliability Engineering (SRE) Architect. In this pivotal role, you will lead the design and implementation of resilient infrastructure and practices to ensure operational efficiency and reliability across services. The ideal candidate will have extensive experience in cloud computing, specifically AWS, along with a strong background in SRE principles and automation tools. This is a long-term position with a focus on developing best practices and mentoring others in the organization.