Especialista de SRE

Jobgether SRL

Brasil

Presencial

BRL 150 000 - 190 000

Tempo integral

há 43 horas
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura completa num minuto — currículo personalizado e carta de apresentação, prontos a enviar.

Ultrapassa os filtros ATS

Resumo da oferta

Jobgether SRL is seeking an experienced Site Reliability Engineer based in Brazil to ensure high availability, performance, and resilience of mission-critical products. You will work with cloud, Kubernetes, IaC, CI/CD, and observability tools to design reliable systems.

As a technical reference for SRE practices, you will define SLIs/SLOs/SLAs, lead incident analysis, automate provisioning, and promote safe change-management across development and operations teams.

Qualificações

  • Solid experience in SRE, DevOps, or Production Engineering within mission-critical environments.
  • Strong expertise with Kubernetes, Docker, and cloud platforms such as AWS, OCI, Azure, and GCP.
  • Advanced knowledge of automation and infrastructure as code, including Terraform and Ansible.
  • Experience with monitoring and observability, particularly Datadog, along with familiarity with Prometheus, ELK, and Grafana.
  • Hands-on experience with CI/CD pipelines, version control, and reliable deployment practices.
  • Strong ability to analyze performance, troubleshoot complex issues, and optimize distributed systems.
  • Knowledge of relational and non-relational databases.
  • Ability to collaborate effectively with development, product, and operations teams.
  • Strong communication, systems thinking, analytical skills, and a problem-solving mindset.
  • Experience with resilience engineering in identity and fraud systems is desirable.
  • Cloud certifications in AWS, OCI, Azure, or GCP are a plus.
  • Experience with chaos engineering and resilience testing is desirable.
  • Knowledge of application and infrastructure security is an advantage.

Responsabilidades

  • Act as the technical SRE reference for Identity & Fraud products, supporting development and operations teams.
  • Define, implement, and monitor SLIs, SLOs, and SLAs aligned with business goals and service expectations.
  • Lead incident analysis and implement preventive and corrective actions to reduce recurring issues.
  • Automate provisioning, deployment, scaling, and failure-recovery processes to improve operational efficiency and reliability.
  • Design and maintain observability solutions covering logs, metrics, traces, and alerts.
  • Support capacity and performance engineering to ensure systems can handle demand predictably.
  • Contribute to architectural improvements focused on resilience, scalability, and security.
  • Promote infrastructure-as-code, CI/CD, version control, and safe change-management practices.
  • Troubleshoot and mitigate issues in real time within critical production environments.

Conhecimentos

Kubernetes
Docker
Cloud platforms AWS
Terraform
Ansible
Datadog
Prometheus
ELK
Grafana
CI/CD
Version control

Ferramentas

Terraform
Ansible
Datadog
Prometheus
ELK
Grafana

Descrição da oferta de emprego

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Especialista de SRE based in Brazil. As a Site Reliability Engineering specialist, you will play a key role in ensuring the reliability and resilience of critical products in a large-scale technology environment. You will serve as a technical reference for SRE practices, partnering closely with development, product, and operations teams. The role focuses on high availability, performance, observability, automation, and secure infrastructure practices. You will help define and monitor SLIs, SLOs, and SLAs aligned with business objectives. Your expertise will contribute to incident prevention, capacity planning, architectural resilience, and efficient recovery from failures. You will work with modern cloud, Kubernetes, infrastructure-as-code, CI/CD, and observability technologies. This is an opportunity to solve complex reliability challenges while strengthening the resilience of mission-critical systems.

Accountabilities
  • Act as the technical SRE reference for Identity & Fraud products, supporting development and operations teams.
  • Define, implement, and monitor SLIs, SLOs, and SLAs aligned with business goals and service expectations.
  • Lead incident analysis and implement preventive and corrective actions to reduce recurring issues.
  • Automate provisioning, deployment, scaling, and failure-recovery processes to improve operational efficiency and reliability.
  • Design and maintain observability solutions covering logs, metrics, traces, and alerts.
  • Support capacity and performance engineering to ensure systems can handle demand predictably.
  • Contribute to architectural improvements focused on resilience, scalability, and security.
  • Promote infrastructure-as-code, CI/CD, version control, and safe change-management practices.
  • Troubleshoot and mitigate issues in real time within critical production environments.
Requirements
  • Solid experience in SRE, DevOps, or Production Engineering within mission-critical environments.
  • Strong expertise with Kubernetes, Docker, and cloud platforms such as AWS, OCI, Azure, and GCP.
  • Advanced knowledge of automation and infrastructure as code, including Terraform and Ansible.
  • Experience with monitoring and observability, particularly Datadog, along with familiarity with Prometheus, ELK, and Grafana.
  • Hands-on experience with CI/CD pipelines, version control, and reliable deployment practices.
  • Strong ability to analyze performance, troubleshoot complex issues, and optimize distributed systems.
  • Knowledge of relational and non-relational databases.
  • Ability to collaborate effectively with development, product, and operations teams.
  • Strong communication, systems thinking, analytical skills, and a problem-solving mindset.
  • Experience with resilience engineering in identity and fraud systems is desirable.
  • Cloud certifications in AWS, OCI, Azure, or GCP are a plus.
  • Experience with chaos engineering and resilience testing is desirable.
  • Knowledge of application and infrastructure security is an advantage.
Benefits
  • Opportunity to work on large-scale, mission-critical technology systems.
  • Collaborative environment involving development, product, and operations teams.
  • People-focused and inclusive workplace culture.
  • Environment designed to support career development, professional growth, and personal well-being.
  • Work-life balance supported alongside career and personal commitments.
  • Opportunities to work with modern cloud, automation, observability, and reliability technologies.
  • Exposure to complex challenges in data, technology, identity, and fraud solutions.
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

SR SRE (Linux & Windows)
SR SRE (Linux & Windows)

Jobgether SRL • Brasil

Presencial
BRL 417 000 - 626 000
Fully remote work
USD-based compensation
Paid time off
+3
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)
SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

iTRTech Group • São Paulo

Presencial
BRL 180 000 - 260 000
Senior Site Reliability / Gitops Engineer
Senior Site Reliability / Gitops Engineer

Jobgether • Brasil

Presencial
BRL 240 000 - 420 000
Annual learning budget
In-person team sprints twice a year
Performance-based bonus or commission
+2
Mid Sre Cloud Product Reliability
Mid Sre Cloud Product Reliability

Serasa Experian • Brasil

Presencial
BRL 180 000 - 300 000
SRE Partner
SRE Partner

Jusbrasil • São Paulo

Presencial
BRL 279 000 - 446 400
SRE Especialista
SRE Especialista

VR • Brasil

Teletrabalho
BRL 180 000 - 240 000
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Presencial
BRL 385 674 - 550 964
Professional growth
Competitive compensation
Flextime
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ernst & Young Advisory Services Sdn Bhd • Pompeia

Presencial
BRL 167 400 - 279 000
DevOps/SRE Engineer - São Paulo, State of São Paulo
DevOps/SRE Engineer - São Paulo, State of São Paulo

MissionHires • Brasil

Presencial
BRL 120 000 - 160 000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Presencial
BRL 298 656 - 398 208
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1