Telemetry SRE: Scale Observability for Open Finance

Kraken

Chile

Presencial

CLP 109.689.000 - 164.534.000

Jornada completa

14 días+
Generador de candidaturas

Transforma esta oferta en una entrevista: un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

Kraken is seeking a Site Reliability Engineer to join the Telemetry team. You will own the shared platform for metrics, logs, traces, alerting, dashboards, and profiling at scale.

Operate and enhance telemetry services with a Prometheus-compatible stack, VictoriaMetrics, Grafana, and modern alerting tools. This role emphasizes distributed systems, automation, and on-call incident response in production environments.

Formación

  • 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
  • Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
  • Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
  • Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
  • Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
  • Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
  • Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
  • Strong incident response, documentation, and collaboration skills.

Responsabilidades

  • Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
  • Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
  • Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
  • Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
  • Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
  • Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
  • Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
  • Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.

Conocimientos

Observability
SRE / Platform engineering
Prometheus
Kubernetes
Terraform
CI/CD
Logging & Telemetry

Herramientas

Prometheus-compatible stack
VictoriaMetrics
Grafana
Tempo
Loki
Vector
Splunk
OpenTelemetry
Pyroscope

Descripción del empleo

Kraken is seeking a Site Reliability Engineer to join the Telemetry team. You will own the shared platform for metrics, logs, traces, alerting, dashboards, and profiling at scale.

Operate and enhance telemetry services with a Prometheus-compatible stack, VictoriaMetrics, Grafana, and modern alerting tools. This role emphasizes distributed systems, automation, and on-call incident response in production environments.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Kubernetes Platform Engineer
Senior Kubernetes Platform Engineer

Kraken • Chile

Presencial
CLP 82.267.000 - 137.112.000
Remote SRE: Build Resilient Systems & Observability
Remote SRE: Build Resilient Systems & Observability

BairesDev • Chile

Presencial
CLP 25.000.000 - 40.000.000
Remote work
Home office setup provided
Flexible hours
+1
Senior SRE: Kubernetes & Observability at Scale (Remote)
Senior SRE: Kubernetes & Observability at Scale (Remote)

BairesDev • Chile

Presencial
CLP 109.689.000 - 146.252.000
Remote work
USD or local currency pay
Home office setup
+4
Senior Rust Engineer – Real-Time FinTech Platform
Senior Rust Engineer – Real-Time FinTech Platform

Kraken • Chile

Presencial
CLP 9.000.000 - 15.000.000
Site Reliability Engineer: Build Resilient, Scalable Cloud
Site Reliability Engineer: Build Resilient, Scalable Cloud

EPAM Systems • Chile

Presencial
CLP 28.000.000 - 48.000.000
Healthcare benefits
Paid time off
LinkedIn Learning access
+2
Senior Observability Engineer - Remote, Flexible Hours
Senior Observability Engineer - Remote, Flexible Hours

BairesDev • Chile

Presencial
CLP 82.267.000 - 118.830.000
100% remote work (from anywhere)
USD or local currency compensation
Hardware and software setup for home
+3
Site Reliability Engineer — Build Resilient, Automated Systems
Site Reliability Engineer — Build Resilient, Automated Systems

Infosys • Santiago

Presencial
CLP 24.000.000 - 42.000.000
Rust Software Engineer — Open-Finance Infrastructure
Rust Software Engineer — Open-Finance Infrastructure

Kraken • Chile

Híbrido
CLP 42.000.000 - 70.000.000
Remote Hybrid SRE for Global E‑commerce Platforms
Remote Hybrid SRE for Global E‑commerce Platforms

Applydigital • Santiago

Híbrido
CLP 56.075.000 - 84.112.000
Generous vacation policy
Flexible work arrangements
AI upskilling & training budgets
+1
OpenStack & Storage Infrastructure Engineer - Scale & Secure
OpenStack & Storage Infrastructure Engineer - Scale & Secure

Kraken • Chile

Presencial
CLP 54.845.000 - 109.689.000