- Review and evolve alerts, monitors, and triggering criteria.
- Ensure proper escalation and reduce ignored or unowned alerts.
- Implement and monitor SLIs, SLOs, and availability metrics.
- Improve platform observability: logs, metrics, tracing, and APM.
- Automate incident responses, diagnostics, and recoveries.
- Create and maintain operational runbooks and playbooks.
- Drive technical actions resulting from post-mortems.
- Reduce recurring failures and operational manual work (toil).
- Support capacity, performance, resilience, and disaster recovery.
- Evolve internal platform components and patterns.
- Provide technical support to DevOps, Platform, and Development teams.
- Explore and apply AIOps solutions for anomaly detection and intelligent alert correlation.
Requirements
- Senior experience in SRE, DevOps, or Platform Engineering.
- Experience with Datadog or an equivalent observability tool.
- Experience with cloud, Kubernetes, and infrastructure as code.
- Knowledge of CI/CD, automation, and scripting or development.
- Experience with incident management and root cause analysis.
- Familiarity with APIs, API gateways, rate limiting, and distributed observability.
- Ability to develop automations and internal tools.
- Hands-on experience with AIOps: anomaly detection, automatic alert correlation, or AI-assisted root cause analysis (plus).
- Experience in high-volume financial or retail environments (plus).
- Cloud (GCP, AWS, or Azure) or Kubernetes (CKA/CKAD) certifications (plus).
Core Competencies
Demonstrates expertise in Site Reliability Engineering (SRE) and DevOps practices, focusing on observability, incident management, and automation. Proficient in cloud technologies and infrastructure as code, with a strong emphasis on improving platform performance and resilience.
Highest-signal resume keywords
- Site Reliability Engineering (SRE)
- Cloud Technologies (GCP, AWS, Azure)
- Kubernetes (CKA/CKAD)
- Incident Management
- AIOps Solutions
ATS Optimization Keywords
Hard Skills
- Observability Tools
- Automation
- Scripting
- CI/CD
- API Management
- Root Cause Analysis
- Metrics Monitoring
- Incident Response
- Performance Optimization
- Disaster Recovery
Certifications & Qualifications
- Cloud Certifications (GCP, AWS, Azure)
- Kubernetes Certifications (CKA, CKAD)
Industry Keywords
- Financial Environments
- Retail Environments
- Operational Runbooks
- Playbooks
- Anomaly Detection
Tools & Technologies
- Datadog
- Kubernetes
- AIOps
- API Gateways
- Monitoring Tools