[Job-32122] Site Reliability Engineer

Lever, Inc.

Brasil

A distancia

BRL 120.000 - 190.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Health and dental insurance
Meal and food allowance
Childcare assistance
Extended paternity leave
WellHub gym partnerships
Profit Sharing and Results PARTICIP.
Life insurance
CI&T University
Discount club
Well-being platform
Pregnancy and parenting course
Online learning partnerships
Language learning platform

Descripción de la vacante

CI&T is hiring a Site Reliability Engineer in Brazil to own the day-to-day operation of our monitoring platform, including dashboards, logs, and APM tools. You will triage incidents, tune alerts, and instrument services to enhance observability.

Ideal candidates have 3+ years in SRE/DevOps, strong troubleshooting skills, and hands-on experience with cloud providers and IaC. Collaboration with security, engineering, and QA is essential for reliable, scalable systems.

Formación

  • 3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused role.
  • Hands-on experience administering and building in platforms such as New Relic, Grafana, Splunk, or Dynatrace including dashboards, monitors, log management, and APM.
  • Strong troubleshooting and root-cause analysis skills, with a demonstrated ability to work through logs, traces, and metrics to find the real problem behind a symptom.
  • Working proficiency in at least one scripting or programming language such as Python, TypeScript/JavaScript, Go, or Bash.
  • Experience operating services in a major cloud provider (AWS, GCP, or Azure), with a solid grasp of networking, containers, and Linux fundamentals.
  • Familiarity with infrastructure as code (Terraform, CloudFormation, or Pulumi) and CI/CD tooling such as GitHub Actions.
  • Experience with on-call responsibilities, incident response, and post-incident review processes.
  • Clear written and verbal communication, including the ability to explain reliability concerns and trade-offs to non-technical stakeholders.

Responsabilidades

  • Own the day-to-day operation of a monitoring platform, including dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring, keeping coverage accurate and current as our services evolve.
  • Proactively sift through logs, error tracking, traces, and metrics to find failures, regressions, and anomalies that have not triggered an alert, then triage, reproduce, and drive them to resolution with the owning team.
  • Tune alert thresholds, monitor logic, and notification routing to reduce noise and false positives while ensuring genuine customer-impacting issues page the right people quickly.
  • Instrument new and existing services with meaningful metrics, structured logs, and distributed traces, and partner with engineers to improve the observability of their code.
  • Write post-incident reviews, track remediation items to completion, and feed lessons learned back into monitors, runbooks, and system design.
  • Define, measure, and report on service level indicators for key customer-facing services.
  • Improve the reliability, scalability, and cost efficiency of our cloud infrastructure, CI/CD pipelines, and release processes, automating repetitive operational work wherever possible.
  • Maintain and improve runbooks, escalation paths, and operational documentation so that the team can act quickly and consistently.
  • Partner with our enterprise InfoSec team to remediate cybersecurity risk items across our infrastructure, including SSL/TLS cleanup, removal of exposed technology and version banners, implementation of security headers (HSTS, CSP, X-Frame-Options, and related), and DNS configuration hygiene (DNSSEC, SPF/DKIM/DMARC, dangling records), tracking findings from scans and audits through to verified closure.
  • Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features before they ship.

Conocimientos

SRE experience
Troubleshooting
Root-cause analysis
Monitoring tools
Scripting languages
Cloud platforms
IaC tooling
CI/CD tooling
On-call
Communication skills

Herramientas

New Relic
Grafana
Splunk
Dynatrace

Descripción del empleo

At CI&T, we help large enterprises transform the potential of AI into real business impact with AI Deployment, AI-native execution, and tech-integrated business solutions.

With 30 years of experience in technological transformation, we accelerate innovation with expertise in Agentic SDLC, Application modernization, Data & AI, Martech and Business strategy.

We are 8,000 CI&Ters across more than 25 countries, collaborating to build solutions with real impact. AI is already part of how we work, evolve, and innovate every day.

Responsibilities
  • Own the day-to-day operation of a monitoring platform, including dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring, keeping coverage accurate and current as our services evolve.
  • Proactively sift through logs, error tracking, traces, and metrics to find failures, regressions, and anomalies that have not triggered an alert, then triage, reproduce, and drive them to resolution with the owning team.
  • Tune alert thresholds, monitor logic, and notification routing to reduce noise and false positives while ensuring genuine customer-impacting issues page the right people quickly.
  • Instrument new and existing services with meaningful metrics, structured logs, and distributed traces, and partner with engineers to improve the observability of their code.
  • Write post-incident reviews, track remediation items to completion, and feed lessons learned back into monitors, runbooks, and system design.
  • Define, measure, and report on service level indicators for key customer-facing services.
  • Improve the reliability, scalability, and cost efficiency of our cloud infrastructure, CI/CD pipelines, and release processes, automating repetitive operational work wherever possible.
  • Maintain and improve runbooks, escalation paths, and operational documentation so that the team can act quickly and consistently.
  • Partner with our enterprise InfoSec team to remediate cybersecurity risk items across our infrastructure, including SSL/TLS cleanup, removal of exposed technology and version banners, implementation of security headers (HSTS, CSP, X-Frame-Options, and related), and DNS configuration hygiene (DNSSEC, SPF/DKIM/DMARC, dangling records), tracking findings from scans and audits through to verified closure.
  • Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features before they ship.
What We're Looking For
  • 3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused software engineering role.
  • Hands-on experience administering and building in platform such as New Relic, Grafana, Splunk, or Dynatrace including dashboards, monitors, log management, and APM.
  • Strong troubleshooting and root-cause analysis skills, with a demonstrated ability to work through logs, traces, and metrics to find the real problem behind a symptom.
  • Working proficiency in at least one scripting or programming language such as Python, TypeScript/JavaScript, Go, or Bash, and comfort reading application code to understand failures.
  • Experience operating services in a major cloud provider (AWS, GCP, or Azure), with a solid grasp of networking, containers, and Linux fundamentals.
  • Familiarity with infrastructure as code (Terraform, CloudFormation, or Pulumi) and CI/CD tooling such as GitHub Actions.
  • Experience with on-call responsibilities, incident response, and post-incident review processes.
  • Clear written and verbal communication, including the ability to explain reliability concerns and trade-offs to non-technical stakeholders.
Our benefits
  • Health and dental insurance
  • Meal and food allowance
  • Childcare assistance
  • Extended paternity leave
  • Partnership with gyms and health and wellness professionals via Wellhub (Gympass) TotalPass;
  • Profit Sharing and Results Participation (PLR);
  • Life insurance
  • Continuous learning platform (CI&T University);
  • Discount club
  • Free online platform dedicated to physical, mental, and overall well-being
  • Pregnancy and responsible parenting course
  • Partnerships with online learning platforms
  • Language learning platform

More details about our benefits here: https://ciandt.com/br/pt-br/carreiras

At CI&T, inclusion starts at the first contact. If you are a person with a disability, it is important to present your assessment during the selection process. See which data needs to be included in the report by clicking here. This way, we can ensure the support and accommodations that you deserve. If you do not yet have the assessment, don't worry: we can support you in obtaining it.

We have a dedicated Health and Well-being team, inclusion specialists, and affinity groups who will be with you at every stage. Count on us to make this journey side by side.

And many more!

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

[Job - 31862] Senior DevOps Engineer, Brazil
[Job - 31862] Senior DevOps Engineer, Brazil

Lever, Inc. • Brasil

Presencial
BRL 240.000 - 420.000
Health insurance
Meal allowance
Wellness program
+1
Senior DevOps Engineer, Brazil
Senior DevOps Engineer, Brazil

CI&T • Brasil

Presencial
BRL 180.000 - 240.000
Health and dental insurance
Meal and food allowance
Extended paternity leave
+2
Senior DevOps Engineer, Brazil
Senior DevOps Engineer, Brazil

Ciandt • Brasil

A distancia
BRL 180.000 - 300.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+1
[Job-32121] Senior Frontend Engineer
[Job-32121] Senior Frontend Engineer

Lever, Inc. • Brasil

Presencial
BRL 120.000 - 180.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+3
Tech Lead (.NET / Angular), Brazil
Tech Lead (.NET / Angular), Brazil

Ciandt • Brasil

A distancia
BRL 180.000 - 300.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+7
[31900] Senior QA Engineer, Brazil
[31900] Senior QA Engineer, Brazil

Ciandt • Brasil

Presencial
BRL 120.000 - 200.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+3
[Job - 32198] Infrastructure and Cloud Engineer, Senior (Brazil)
[Job - 32198] Infrastructure and Cloud Engineer, Senior (Brazil)

Lever, Inc. • Brasil

Presencial
BRL 180.000 - 280.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+4
Senior Data Developer (Analytics Engineer), Brazil
Senior Data Developer (Analytics Engineer), Brazil

CI&T • Brasil

Presencial
BRL 180.000 - 290.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+6
- Software Architect, Brazil
- Software Architect, Brazil

CI&T • Brasil

Presencial
BRL 260.000 - 420.000
Health and dental insurance
Meal and food allowance
Extended paternity leave
+3
Senior Data Developer/AI Solutions Developer, Brazil
Senior Data Developer/AI Solutions Developer, Brazil

Ciandt • Brasil

A distancia
BRL 110.000 - 180.000
Health and dental insurance
Meal and food allowance
Childcare assistance
+2