Destaca en este puesto — genera un currículum adaptado y una carta de presentación en aproximadamente un minuto.
CLOUDSUFI is seeking a Lead Site Reliability Engineer to drive enterprise reliability across production environments. You will own incident response, postmortems, and RCAs while shaping SLI/SLO frameworks and error budgets.
The role emphasizes building robust observability, dashboards, and alerts, plus automation via Python/Bash and API integrations. Mentorship of SRE/DevOps peers and collaboration with cross-functional teams are essential components.
Hands-on experience with Datadog or an equivalent APM, logging, and monitoring platform such as New Relic, Dynatrace, Grafana/Prometheus, Splunk, ELK/OpenSearch, CloudWatch, or AppDynamics.
Experience should include:
Hands-on experience with PagerDuty or an equivalent alerting/paging platform, such as Opsgenie, ServiceNow, Splunk On-Call, xMatters, or Grafana Alerting.