Telemetry SRE: Scale Observability & Incidents

Kraken

Argentina

Presencial

ARS 163.756.000 - 253.078.000

Jornada completa

14 días+
Generador de candidaturas

Destaca para este puesto: genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Descripción de la vacante

Kraken is seeking a Site Reliability Engineer to join the Telemetry team, focusing on the shared platform for metrics, logs, traces, and profiling. You will maintain and improve Prometheus-compatible stacks, log pipelines, and distributed tracing across environments, with IaC, container orchestration, and on-call responsibilities.

Ideal candidates have 3+ years in production engineering, strong incident response skills, and a hands-on approach to observability, performance, and scale in a

Formación

  • 3+ years in a Site Reliability Engineer, Platform/Infrastructure/Observability role.
  • Experience with Prometheus or Prometheus-compatible monitoring stacks.
  • Experience troubleshooting distributed production systems.
  • IaC with Terraform and CI/CD.
  • Containerised workloads with Nomad, Kubernetes, or similar.
  • Strong scripting/programming ability and use of AI tools to accelerate delivery.
  • Strong incident response, documentation, and collaboration skills.

Responsabilidades

  • Operate and improve the shared telemetry platform for metrics, logs, traces, alerting, dashboards, and profiling.
  • Maintain metrics collection, long-term storage, querying, dashboards, and alerting.
  • Operate log pipelines using Vector, Splunk, and Loki; ensure reliability and throughput.
  • Handle distributed tracing and profiling with Grafana Alloy, Tempo, OpenTelemetry, Pyroscope.
  • Deploy telemetry services using Terraform, Terragrunt, and container orchestration.
  • Troubleshoot missing data, slow queries, broken alerts, and capacity issues.
  • Build reusable configurations and automation for dashboards, alerts, and telemetry integrations.
  • Participate in incident response and on-call; write runbooks and improve platform from incidents.

Conocimientos

SRE
Observability
Incident response
Automation

Herramientas

Terraform
Terragrunt
Nomad
Kubernetes
Grafana
Tempo
OpenTelemetry
Pyroscope
Vector
Splunk

Descripción del empleo

Kraken is seeking a Site Reliability Engineer to join the Telemetry team, focusing on the shared platform for metrics, logs, traces, and profiling. You will maintain and improve Prometheus-compatible stacks, log pipelines, and distributed tracing across environments, with IaC, container orchestration, and on-call responsibilities.

Ideal candidates have 3+ years in production engineering, strong incident response skills, and a hands-on approach to observability, performance, and scale in a

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Hybrid SRE Engineer — Observability & Incident Response
Hybrid SRE Engineer — Observability & Incident Response

APPLY • Comisión de Fomento de Perú

Híbrido
ARS 105.910.000 - 166.430.000
Agentic Delivery
Inclusive culture
AI upskilling budget
+2
Site Reliability Engineer - Telemetry
Site Reliability Engineer - Telemetry

Kraken • Argentina

Presencial
ARS 163.756.000 - 253.078.000
Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 3.500.000 - 6.000.000
Remote SRE for Global E-commerce Reliability
Remote SRE for Global E-commerce Reliability

APPLY • Argentina

Híbrido
ARS 75.496.000 - 105.694.000
Agentic Delivery
Inclusive culture
Training budget
+2
Senior Azure SRE: Reliability, Observability & Scale
Senior Azure SRE: Reliability, Observability & Scale

Teladoc Health, Inc. • Argentina

A distancia
ARS 166.279.000 - 226.744.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Teladoc Health • Argentina

Presencial
ARS 136.170.000 - 196.690.000
Senior SRE — Remote, Multi-Cloud & Incident Commander
Senior SRE — Remote, Multi-Cloud & Incident Commander

AgileEngine • Argentina

Híbrido
ARS 2.000.000 - 4.000.000
Professional growth
Competitive USD-based compensation
Flextime
+1
Senior SRE & Platform Engineer: Observability & Automation
Senior SRE & Platform Engineer: Observability & Automation

Techunting • Córdoba

Presencial
ARS 3.500.000 - 6.000.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off
Upskilling programs
+2
Senior SRE: Build Resilient Cloud, Automate & Scale
Senior SRE: Build Resilient Cloud, Automate & Scale

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off
Upskilling programs
+2