Stand out for this role — generate a tailored resume and cover letter in about a minute.
Acldigital is looking for an experienced Site Reliability Engineer with 8+ years in large-scale enterprise environments. The role emphasizes strong expertise in AWS, Kubernetes administration, and comprehensive incident management using tools like PagerDuty.
Ideal candidates should have a solid background in automation, observability tools such as Datadog and Grafana, and the ability to enhance system reliability. Continuous improvement and proactive monitoring are key responsibilities in this position.
8+ years of experience in Site Reliability Engineering(SRE)
Experience in large-scale enterprise environments
Strong expertise in AWS
Terraform or similar IaC tools
Hands-on experience with Datadog, Grafana, and Prometheus
Advanced cloud optimization practices
Experience with PagerDuty for incident management
Kubernetes administration and troubleshooting
Security and compliance automation
Automation and scripting expertise
Advanced observability practices
Infrastructure as Code (IaC) implementation