Senior Reliability and Platform Engineer

Jobtailor

São Paulo

Presencial

BRL 180 000 - 300 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Jobtailor seeks a Senior Site Reliability Engineer to advance observability, automate incident response, and improve platform resilience across cloud and Kubernetes environments.

You will work with Datadog, CI/CD pipelines, and AIOps approaches to reduce toil and enhance disaster recovery, while building runbooks and playbooks for operations.

Qualificações

  • Senior experience in SRE, DevOps, or Platform Engineering.
  • Experience with Datadog or an equivalent observability tool.
  • Experience with cloud, Kubernetes, and infrastructure as code.
  • Knowledge of CI/CD, automation, and scripting or development.
  • Experience with incident management and root cause analysis.
  • Familiarity with APIs, API gateways, rate limiting, and distributed observability.
  • Ability to develop automations and internal tools.
  • Hands-on experience with AIOps: anomaly detection, automatic alert correlation, or AI-assisted root cause analysis (plus).
  • Experience in high-volume financial or retail environments (plus).
  • Cloud Certifications (GCP, AWS, or Azure) or Kubernetes Certifications (CKA, CKAD) (plus).

Responsabilidades

  • Review and evolve alerts, monitors, and triggering criteria.
  • Ensure proper escalation and reduce ignored or unowned alerts.
  • Implement and monitor SLIs, SLOs, and availability metrics.
  • Improve platform observability: logs, metrics, tracing, and APM.
  • Automate incident responses, diagnostics, and recoveries.
  • Create and maintain operational runbooks and playbooks.
  • Drive technical actions resulting from post-mortems.
  • Reduce recurring failures and operational manual work (toil).
  • Support capacity, performance, resilience, and disaster recovery.
  • Evolve internal platform components and patterns.
  • Provide technical support to DevOps, Platform, and Development teams.
  • Explore and apply AIOps solutions for anomaly detection and intelligent alert correlation.

Conhecimentos

Site Reliability Engineering (SRE)
Cloud Technologies (GCP AWS Azure)
Kubernetes (CKA CKAD)
Incident Management
AIOps Solutions
Observability
Automation
Scripting
CI/CD
API Management

Formação académica

Cloud Certification (GCP/AWS/Azure)
Kubernetes Certification (CKA/CKAD)

Ferramentas

Datadog
Kubernetes
AIOps
API Gateways
Monitoring Tools

Descrição da oferta de emprego

  • Review and evolve alerts, monitors, and triggering criteria.
  • Ensure proper escalation and reduce ignored or unowned alerts.
  • Implement and monitor SLIs, SLOs, and availability metrics.
  • Improve platform observability: logs, metrics, tracing, and APM.
  • Automate incident responses, diagnostics, and recoveries.
  • Create and maintain operational runbooks and playbooks.
  • Drive technical actions resulting from post-mortems.
  • Reduce recurring failures and operational manual work (toil).
  • Support capacity, performance, resilience, and disaster recovery.
  • Evolve internal platform components and patterns.
  • Provide technical support to DevOps, Platform, and Development teams.
  • Explore and apply AIOps solutions for anomaly detection and intelligent alert correlation.
Requirements
  • Senior experience in SRE, DevOps, or Platform Engineering.
  • Experience with Datadog or an equivalent observability tool.
  • Experience with cloud, Kubernetes, and infrastructure as code.
  • Knowledge of CI/CD, automation, and scripting or development.
  • Experience with incident management and root cause analysis.
  • Familiarity with APIs, API gateways, rate limiting, and distributed observability.
  • Ability to develop automations and internal tools.
  • Hands-on experience with AIOps: anomaly detection, automatic alert correlation, or AI-assisted root cause analysis (plus).
  • Experience in high-volume financial or retail environments (plus).
  • Cloud (GCP, AWS, or Azure) or Kubernetes (CKA/CKAD) certifications (plus).
Core Competencies

Demonstrates expertise in Site Reliability Engineering (SRE) and DevOps practices, focusing on observability, incident management, and automation. Proficient in cloud technologies and infrastructure as code, with a strong emphasis on improving platform performance and resilience.

Highest-signal resume keywords
  • Site Reliability Engineering (SRE)
  • Cloud Technologies (GCP, AWS, Azure)
  • Kubernetes (CKA/CKAD)
  • Incident Management
  • AIOps Solutions
ATS Optimization Keywords
Hard Skills
  • Observability Tools
  • Automation
  • Scripting
  • CI/CD
  • API Management
  • Root Cause Analysis
  • Metrics Monitoring
  • Incident Response
  • Performance Optimization
  • Disaster Recovery
Certifications & Qualifications
  • Cloud Certifications (GCP, AWS, Azure)
  • Kubernetes Certifications (CKA, CKAD)
Industry Keywords
  • Financial Environments
  • Retail Environments
  • Operational Runbooks
  • Playbooks
  • Anomaly Detection
Tools & Technologies
  • Datadog
  • Kubernetes
  • AIOps
  • API Gateways
  • Monitoring Tools
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Mid-level SRE
Mid-level SRE

Jobtailor • São Paulo

Presencial
BRL 180 000 - 240 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 260 000
Senior DevOps Analyst – SRE, Kubernetes, Cloud
Senior DevOps Analyst – SRE, Kubernetes, Cloud

Jobtailor • São Paulo

Presencial
BRL 90 000 - 130 000
DevOps Engineer I
DevOps Engineer I

Jobtailor • Blumenau

Presencial
BRL 90 000 - 150 000
Technology Intern – Operations and SRE
Technology Intern – Operations and SRE

Jobtailor • São Paulo

Híbrido
BRL 13 000 - 28 000
Senior Infrastructure and Cloud Analyst
Senior Infrastructure and Cloud Analyst

Jobtailor • São Paulo

Presencial
BRL 180 000 - 240 000
DevOps Coordinator – SRE
DevOps Coordinator – SRE

Jobtailor • São Paulo

Presencial
BRL 180 000 - 320 000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Brasil

Presencial
BRL 250 000 - 450 000
Technical Monitoring Analyst, Junior
Technical Monitoring Analyst, Junior

Jobtailor • São Paulo

Presencial
BRL 60 000 - 100 000
Senior Software Engineer – C++, Python
Senior Software Engineer – C++, Python

Jobtailor • São Paulo

Presencial
BRL 150 000 - 210 000