Expert Observability Engineer

Ensono

São Paulo

Presencial

BRL 350 000 - 550 000

Tempo integral

há 42 horas
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura completa num minuto — currículo e carta de apresentação personalizados, prontos a enviar.

Ultrapassa os filtros ATS

Resumo da oferta

Ensono seeks a senior observability leader to architect and govern enterprise observability across multi-cloud environments, using IBM Instana, Grafana, Prometheus, and OpenTelemetry. You will lead FOAK implementations and set standards for telemetry pipelines and cost management.

You will drive SRE practices, define SLIs/SLOs, reduce MTTR, and mentor global teams in a 24x7 setting, with a focus on security and RBAC.

Qualificações

  • 12+ years of total IT experience.
  • 5 to 7 years functioning as Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise.
  • Proven track record migrating from legacy monitoring to proactive observability platforms.

Responsabilidades

  • Architect and govern a unified observability framework covering metrics, logs, traces, and events.
  • Lead FOAK implementations—evaluating new observability tech and converting them into secure, repeatable patterns.
  • Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.
  • Define and govern SLIs, Objectives (SLOs), and error budgets.
  • Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting RCA.

Conhecimentos

IBM Instana
Grafana
Prometheus
OpenTelemetry
Telegraf
InfluxDB
SolarWinds
Netcool
Elastic
Splunk
Linux
Windows Server
VMware
Citrix VDI
Load Balancers
Edge Proxies
Ansible
Terraform
Python
Bash
Webhooks
GitHub Actions
GitLab
Jenkins
ServiceNow
ITIL 4
Major Incident Mgmt

Ferramentas

IBM Instana
Grafana
Prometheus
OpenTelemetry
Telegraf
InfluxDB
SolarWinds
Netcool
Elastic
Splunk
Linux
Windows Server
VMware
Citrix VDI
Load Balancers
Edge Proxies
Ansible
Terraform
Python
Bash
Webhooks
GitHub Actions
GitLab
Jenkins
ServiceNow
ITIL 4
Major Incident Mgmt

Descrição da oferta de emprego

At Ensono, our Purpose is to be a relentless ally, disrupting the status quo and unleashing our clients to Do Great Things! We enable our clients to achieve key business outcomes that reshape how our world runs. As an expert technology adviser and managed service provider with cross-platform certifications, Ensono empowers our clients to keep up with continuous change and embrace innovation.

We can Do Great Things because we have great Associates. The Ensono Core Values unify our diverse talents and are woven into how we do business. These five traits are the key to achieving our purpose.

Job Responsibilities:
1. Enterprise Architecture & Strategy
  • Architect and govern a unified observability framework covering metrics, logs, traces, and events using IBM Instana, Grafana, OpenTelemetry, Telegraf, and InfluxDB.
  • Lead First-of-a-Kind (FOAK) implementations—evaluating new observability tech and converting them into secure, repeatable, production-ready patterns.
  • Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.
2. SRE & Service Reliability
  • Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets.
  • Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting evidence-based Root Cause Analysis (RCA).
  • Drastically reduce MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping.
  • Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps.
  • Automate the deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors.
  • Integrate observability platforms seamlessly with ITSM (ServiceNow), Netcool, and CI/CD pipelines.
  • Design deep observability for Docker, Kubernetes, microservices, and multi-cloud environments (Azure/AWS/GCP).
  • Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.
  • Ensure secure-by-design telemetry pipelines (RBAC, TLS, secrets management, and image scanning).
5. Technical Leadership & Transition Management
  • Lead complex Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers for global 24x7 teams.
  • Mentor cross-functional engineering teams and influence enterprise technology roadmaps.
Required Qualifications:
  • Observability & APM: IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, InfluxDB.
  • Legacy/Traditional Monitoring: SolarWinds, Netcool, Elastic/Splunk.
  • Infrastructure: Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
  • Automation & DevOps: Ansible, Terraform, Python, Bash, Webhooks, CI/CD (GitHub Actions/GitLab/Jenkins).
  • ITSM/Operations: ServiceNow, ITIL 4, advanced Major Incident Management.
Qualifications & Experience
  • 12+ years of total IT experience, with a minimum of 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise environment.
  • Proven track record of migrating organizations from legacy monitoring to proactive, automated observability platforms.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Lead Operational Intelligence Engineer
Lead Operational Intelligence Engineer

EPAM Systems • Brasil

Presencial
BRL 300 000 - 420 000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Recife

Híbrido
BRL 298 000 - 399 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior Observability Architect | West Coast | PST | Remote
Senior Observability Architect | West Coast | PST | Remote

DevOpsChat • Costa Rica

Híbrido
BRL 777 000 - 984 000
Observability Engineer – PL/SR
Observability Engineer – PL/SR

Stefanini Brasil • Brasil

Teletrabalho
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Florianópolis

Híbrido
BRL 414 000 - 622 000
Professional growth
Competitive compensation
Exciting projects
+1
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • São Bernardo do Campo

Híbrido
BRL 517 000 - 673 000
Professional growth
Competitive compensation
Exciting projects
+1
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Recife

Híbrido
BRL 180 000 - 250 000
Professional growth
Competitive compensation
Exciting projects
+1
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Brasília

Híbrido
BRL 414 000 - 622 000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Riograndina

Híbrido
BRL 385 000 - 551 000
Professional growth
Competitive compensation
Flextime
+1
Service Delivery Lead
Service Delivery Lead

Espire Infolabs (Singapore) Pte Ltd • Região Norte

Híbrido
BRL 150 000 - 260 000