SRE Consultant – Site Reliability Engineering

Jobtailor

São Paulo

Presencial

BRL 180 000 - 320 000

Tempo integral

Há 3 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura completa num minuto — currículo e carta de apresentação personalizados, prontos a enviar.

Ultrapassa os filtros ATS

Resumo da oferta

Jobtailor in São Paulo, Brazil, seeks an experienced Site Reliability Engineer to own the reliability of Orders, Portability, and Digital Services systems. You will define and monitor SLOs, SLIs, and error budgets; lead proactive automation, runbooks, and AI agents for incident prevention and rapid remediation.

Join a team that collaborates with São Paulo and Curitiba squads to improve deployments, observability, and resilience while reducing toil.

Qualificações

  • 4+ years of experience in SRE/DevOps/Platform Engineering.
  • Experience with cloud environments: AWS, Azure, or GCP.
  • Proficiency in Kubernetes, Docker, Terraform, and CI/CD.
  • Knowledge of Jenkins, GitLab CI, and Azure DevOps.
  • Knowledge of Java 17/21, Spring Boot, REST APIs, and JDBC.
  • Knowledge of Oracle, OCI, databases, and application servers.
  • Knowledge of Angular.
  • Knowledge of microservices, Postman, REST, and caching.
  • Observability using Prometheus, Grafana, ELK, Datadog, or similar tools.
  • Experience in incident management, on-call rotations, and postmortems.
  • Knowledge of SLOs, SLIs, and Error Budgets.
  • Experience in telecommunications companies is a plus.
  • Knowledge of Orders, Portability, and Digital Services systems is a plus.
  • Experience with AI agents and proactive automation is a plus.
  • Data-driven mindset with a results-oriented approach.
  • Clear communication skills for managing crises and engaging stakeholders.
  • Proactive approach to identifying bottlenecks before they become incidents.
  • Collaborative work with teams in São Paulo and Curitiba.

Responsabilidades

  • Own the reliability of Orders, Portability, and Digital Services systems.
  • Define and monitor SLOs, SLIs, and Error Budgets.
  • Work to meet SLAs and reduce MTTR/MTTD.
  • Build proactive automations, runbooks, and AI agents for incident prevention and self-remediation.
  • Enhance the metrics, logging, and tracing stack.
  • Ensure actionable alerts and system health dashboards.
  • Participate in the on-call rotation.
  • Lead postmortems with root-cause analysis (RCA) and action plans.
  • Support development teams in designing cloud-native applications, DR, chaos engineering, and load testing.
  • Improve deployment pipelines and promote SRE best practices alongside development teams.
  • Foster an SRE, blameless, and continuous improvement culture.
  • Provide technical leadership to the area’s squads and reduce toil while increasing system resilience and predictability.

Conhecimentos

SRE / DevOps mindset
Kubernetes
Docker
Terraform
CI/CD pipelines
Observability tools
Incident management
On-call experience

Ferramentas

Jenkins
GitLab CI
Azure DevOps
Postman

Descrição da oferta de emprego

  • Serve as the owner of the reliability of Orders, Portability, and Digital Services systems
  • Define and monitor SLOs, SLIs, and Error Budgets
  • Work to meet SLAs and reduce MTTR/MTTD
  • Build proactive automations, runbooks, and AI agents for incident prevention and self-remediation
  • Enhance the metrics, logging, and tracing stack
  • Ensure actionable alerts and system health dashboards
  • Participate in the on-call rotation
  • Lead postmortems with root-cause analysis (RCA) and action plans
  • Support development teams in designing cloud-native applications, disaster recovery (DR), chaos engineering, and load testing
  • Improve deployment pipelines and promote SRE best practices alongside development teams
  • Foster an SRE, blameless, and continuous improvement culture
  • Provide technical leadership to the area’s squads and reduce toil while increasing system resilience and predictability
Requirements
  • 4+ years of experience in SRE, DevOps, or Platform Engineering
  • Experience with cloud environments: AWS, Azure, or GCP
  • Proficiency in Kubernetes, Docker, Terraform, and CI/CD
  • Knowledge of Jenkins, GitLab CI, and Azure DevOps
  • Knowledge of Java 17/21, Spring Boot, REST APIs, and JDBC
  • Knowledge of Oracle, OCI, databases, and application servers
  • Knowledge of Angular
  • Knowledge of microservices, Postman, REST, and caching
  • Knowledge of observability using Prometheus, Grafana, ELK, Datadog, or similar tools
  • Experience with incident management, on-call rotations, and postmortems
  • Knowledge of SLOs, SLIs, and Error Budgets
  • Experience in telecommunications companies is a plus
  • Knowledge of Orders, Portability, and Digital Services systems is a plus
  • Experience with AI agents and proactive automation is a plus
  • Data-driven mindset with a results-oriented approach
  • Clear communication skills for managing crises and engaging stakeholders
  • Proactive approach to identifying bottlenecks before they become incidents
  • Collaborative work with teams in São Paulo and Curitiba
Core Competencies

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on cloud-native application design, incident management, and proactive automation. Capable of defining and monitoring SLOs, SLIs, and Error Budgets while fostering a culture of continuous improvement and collaboration.

Highest-signal resume keywords
  • Site Reliability Engineering (SRE)
  • Cloud Environments: AWS, Azure, GCP
  • Kubernetes, Docker, Terraform
  • Incident Management and Postmortems
  • Proactive Automation and AI Agents
ATS Optimization Keywords
Hard Skills
  • SLOs, SLIs, and Error Budgets
  • Java 17/21, Spring Boot
  • REST APIs and JDBC
  • Microservices and Caching
  • Observability: Prometheus, Grafana, ELK, Datadog
Soft Skills
  • Clear Communication Skills
  • Proactive Approach
  • Collaborative Work
Industry Keywords
  • Telecommunications
  • Orders, Portability, and Digital Services Systems
Tools & Technologies
  • Jenkins
  • GitLab CI
  • Azure DevOps
  • Postman
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP SE • São Leopoldo

Presencial
BRL 201 000 - 312 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Senior SRE Engineer
Senior SRE Engineer

Jobtailor • São Paulo

Presencial
BRL 156 000 - 257 000
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP SE • Brasil

Presencial
BRL 180 000 - 320 000
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP • São Leopoldo

Presencial
BRL 180 000 - 260 000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Proative Technology • Barueri

Presencial
BRL 180 000 - 240 000
Service Delivery Lead
Service Delivery Lead

Espire Infolabs (Singapore) Pte Ltd • Região Norte

Híbrido
BRL 150 000 - 260 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Brasil

Híbrido
BRL 180 000 - 360 000
Flexible work format
Competitive salary
Career growth
+3