SRE Pleno

Jobgether

Brasil

Presencial

BRL 120 000 - 210 000

Tempo integral

Há 5 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Destaca-te para esta função — gera um currículo e uma carta de apresentação personalizados em cerca de um minuto.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Observability program
Multi-region exposure
OpenTelemetry stack
SRE framework adoption
Learning culture

Resumo da oferta

Jobgether is seeking a Pleno SRE based in Brazil to strengthen the reliability and resilience of production services across multiple regions. You will contribute to a strategic observability initiative, developing SRE practices and maturity alongside technical leadership.

The role blends hands-on troubleshooting with proactive reliability engineering on AWS and on‑prem environments. You will focus on automation, incident management, scalability, and FinOps awareness for cost efficiency.

Qualificações

  • Solid production experience with distributed systems and incident scenarios.
  • Experience with AWS production environments (EC2, networking, IAM) and on-prem may be advantageous.
  • Hands-on with container tech and orchestration for production workloads.
  • Familiarity with observability tooling (Grafana, Prometheus, OpenTelemetry).
  • Understanding of SRE concepts: SLIs/SLOs, error budgets, postmortems, and incident response.
  • Strong Linux networking and web protocols knowledge.
  • DevOps practices including CI/CD and IaC are desirable.

Responsabilidades

  • Identify operational risks and noisy alerts to prevent incidents.
  • Contribute to observability platform evolution across logs, metrics, and tracing.
  • Lead root-cause analyses and postmortems with corrective action tracking.
  • Respond to and troubleshoot critical production incidents across AWS and on‑prem.
  • Identify FinOps opportunities and drive cost efficiency in cloud infra.
  • Collaborate with developers to improve reliability, performance, and readiness.
  • Support scalability initiatives with automation and repeatable processes.
  • Help mature the team’s SRE practices and incident culture.

Conhecimentos

AWS
Docker
Grafana
Prometheus
OpenTelemetry
Terraform
CI/CD

Ferramentas

Kubernetes
Python
Bash
Ansible

Descrição da oferta de emprego

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Pleno based in Brazil.


This role offers the opportunity to contribute directly to the reliability and resilience of production services supporting customers across multiple regions worldwide. You will be part of a strategic observability initiative, helping build SRE processes, practices, and maturity alongside technical leadership. The position combines hands-on troubleshooting with preventive reliability engineering across AWS cloud and on-premises environments. You will work to identify operational risks before they become incidents and turn root causes into actionable improvements. Your responsibilities will span observability, incident management, infrastructure automation, scalability, and cloud cost awareness. This is an environment where autonomy, collaboration, resilience, and a proactive approach to reliability are highly valued.


Accountabilities


  • Proactively identify operational risks, potential points of failure, and noisy alerts in production environments, implementing improvements that prevent incidents and reduce recurring reliability issues.

  • Contribute to the development and evolution of the observability platform across logs, metrics, and distributed tracing, helping define and implement SLIs, SLOs, and error budgets.

  • Conduct root-cause analyses and lead postmortems following significant incidents, documenting technical and operational learnings and tracking corrective actions through completion.

  • Respond to and troubleshoot critical production incidents across AWS and on-premises environments, helping restore services quickly while maintaining clear communication throughout the incident lifecycle.

  • Identify FinOps opportunities and contribute to greater predictability and efficiency in cloud infrastructure costs.

  • Collaborate closely with development teams to improve application reliability, performance, resilience, and operational readiness throughout the software lifecycle.

  • Support scalability initiatives and the creation of new infrastructure, with a strong focus on automation, repeatability, and operational efficiency.

  • Contribute to the evolution of the team’s SRE maturity by strengthening processes, documentation, incident practices, reliability standards, and a culture of continuous learning.


Requirements


  • Solid professional experience troubleshooting production environments and distributed systems, with the ability to diagnose complex operational and performance issues under real-world conditions.

  • Hands-on experience working with AWS in production environments, particularly EC2, networking, load balancing, and IAM, with experience in on-premises environments considered an advantage.

  • Practical experience with Docker and Docker Compose, including the operation and troubleshooting of containerized applications in production.

  • Experience with observability and monitoring platforms such as Grafana, Prometheus, Datadog, SigNoz, or similar technologies, together with knowledge of OpenTelemetry for logs, metrics, and distributed tracing.

  • Strong understanding of SRE principles and practices, including SLI/SLO definition, error budgets, incident management, troubleshooting, root-cause analysis, and postmortems.

  • Solid knowledge of Linux, networking concepts, and protocols such as HTTP, TCP/IP, and DNS.

  • Strong communication skills, autonomy, and resilience, with the ability to remain effective during critical incidents and, when necessary, communicate directly with customers.

  • Experience with automation technologies such as Python, Bash, Terraform, Ansible, or similar tools is considered a strong advantage.

  • Familiarity with DevOps practices, including CI/CD pipelines and Infrastructure as Code (IaC), is desirable.


Benefits


  • Opportunity to work on a strategic observability initiative and contribute directly to the evolution of SRE maturity and reliability practices.

  • Hands-on experience across AWS cloud and on-premises production environments supporting services used by customers in multiple regions.

  • Exposure to modern observability technologies, including logs, metrics, tracing, OpenTelemetry, and SRE reliability frameworks.

  • Opportunity to develop expertise in incident management, postmortems, automation, scalability, infrastructure, and FinOps practices.

  • Close collaboration with technical leadership and development teams, providing opportunities to influence engineering and reliability practices.

  • Strong environment for continuous learning, technical growth, and practical application of modern SRE and DevOps methodologies.

  • Opportunity to make a direct impact on production reliability, operational efficiency, and customer experience.


We appreciate your interest and wish you the best!


Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Híbrido
BRL 180 000 - 260 000
Engenheiro SRE Especialista
Engenheiro SRE Especialista

Jobgether • Brasil

Presencial
BRL 180 000 - 240 000
Mission-critical AWS DBs
Cloud-native tooling
Professional development
+1
Senior Site Reliability / Gitops Engineer
Senior Site Reliability / Gitops Engineer

Jobgether • Brasil

Presencial
BRL 240 000 - 420 000
Annual learning budget
In-person team sprints twice a year
Performance-based bonus or commission
+2
DevOps/SRE Engineer - São Paulo, State of São Paulo
DevOps/SRE Engineer - São Paulo, State of São Paulo

MissionHires • Brasil

Presencial
BRL 120 000 - 160 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Analista SRE Sênior (AWS | OCI) - PRESENCIAL (PRESENCIAL - SÃO PAULO-SP)
Analista SRE Sênior (AWS | OCI) - PRESENCIAL (PRESENCIAL - SÃO PAULO-SP)

Cedro Technologies • São Paulo

Presencial
BRL 180 000 - 280 000
15 dias de descanso remunerados após 1
Day off no aniversário
Site Reliability Engineer - SAP Cloud Ops
Site Reliability Engineer - SAP Cloud Ops

SAP SE • São Leopoldo

Presencial
BRL 201 000 - 312 000
Site Reliability Engineer ID55632
Site Reliability Engineer ID55632

AgileEngine • São Paulo

Híbrido
Professional growth opportunities
Competitive USD-based compensation
Exciting projects with top companies
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Brasil

Híbrido
BRL 180 000 - 360 000
Flexible work format
Competitive salary
Career growth
+3