Platform Observability Engineer

StoneX Group Inc.

Bogotá

Híbrido

COP 310.578.000 - 465.867.000

Jornada completa

hace 12 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Medical and life insurance
Public Transportation Support
Meal and food allowances

Descripción de la vacante

StoneX Group Inc. in São Paulo and Bogota seeks a Platform Observability Engineer to design, build, and operate the observability tooling stack. You’ll own OpenSearch logging on Kubernetes, extend with Datadog, and drive telemetry via OpenTelemetry.

Collaboration with security, development, and operations will improve reliability and MTTR. This role offers a hybrid model with four days in the office and one remote in a dynamic, global financial services environment.

Formación

  • A track record of self-driven problem solving with minimal oversight.
  • 5+ years in observability, SRE, platform, or infrastructure roles.
  • Hands-on experience with Datadog or similar platforms (metrics, logs, APM, tracing, dashboards, alerting).
  • Experience operating OpenSearch/Elasticsearch on Kubernetes.
  • Proficiency with OpenTelemetry and data pipelines (Cribl, Fluent Bit, Logstash, Kafka).
  • Terraform, Docker, Kubernetes, Git, Helm in teams.

Responsabilidades

  • Design, implement, and operate scalable observability platforms with focus on OpenSearch and Kubernetes.
  • Operate and improve OpenSearch clusters, manage index lifecycle, diagnose performance issues.
  • Drive telemetry by default across services and platforms using OpenTelemetry.
  • Build data pipelines and dashboards to reduce MTTR and improve reliability.
  • Collaborate with security, compliance, architecture, and development teams.

Conocimientos

Observability engineering
SRE
Cloud infrastructure
Python/Go scripting
Linux fundamentals
Security best practices
Collaboration

Educación

Bachelor’s degree in computer science, engineering, or related field

Herramientas

Datadog
Prometheus
Grafana
OpenSearch/Elasticsearch
Kubernetes
Terraform
Docker
Ansible
OpenTelemetry
Cribl
Fluent Bit
Kafka

Descripción del empleo

Overview

This role can be located in Sao Paulo (Brazil) or Bogota (Colombia)

Please submit CV in English

Connecting clients to markets – and talent to opportunity. With 4,300 employees and over 400,000 retail and institutional clients from more than 80 offices spread across five continents, we’re a Fortune-100, Nasdaq-listed provider, connecting clients to the global markets – focusing on innovation, human connection, and providing world-class products and services to all types of investors.

Corporate: Engage in a deep variety of business-critical activities that keep our company running efficiently. From strategic marketing and financial management to human resources and operational oversight, you’ll have the opportunity to optimize processes and implement game-changing policies.

Responsibilities

Job Purpose: As a Platform Observability Engineer, you will design, build, and operate the platforms, tooling, and practices that provide end to end visibility into our systems. In the near term your primary focus will be owning and evolving our OpenSearch based logging and search platform running on Kubernetes. Over time you will work more broadly across our observability stack centered on Datadog or a similar platform, integrating complementary open source and cloud native technologies to meet strategic business goals. This role requires expertise in metrics, logs, traces, alerting, SLOs, and instrumentation, strong skills in infrastructure as code and automation, and close collaboration with security, development, product, and operations teams. You will be a hands‑on contributor who improves reliability, accelerates troubleshooting, and enables data driven engineering through first class observability.

Primary duties will include:
  • Design, implement, and operate scalable observability platforms and services, with a primary focus on our OpenSearch based logging and search platform on Kubernetes, and Datadog or a similar platform
  • Operate and improve OpenSearch clusters in production, including scaling, upgrades, index lifecycle management, and troubleshooting performance and reliability issues
  • Adopt instrumentation standards using OpenTelemetry where appropriate, drive telemetry by default across services and platforms
  • Design, build, and manage observability data pipelines that extract, transform, and route telemetry from diverse sources using tools like Cribl, Vector, Fluent Bit, or OpenTelemetry Collector
  • Implement data filtering, enrichment, redaction, and normalization to ensure telemetry quality, compliance, and cost efficiency
  • Optimize observability data lifecycle management, including tiered storage, retention policies, and archive strategies
  • Build actionable alerting, dashboards, runbooks, and analytics that reduce noise and improve MTTR
  • Implement infrastructure as code for observability resources using Terraform, including Datadog or similar providers and reusable modules
  • Integrate observability into CI and CD workflows, enable self‑service patterns for developers and SREs
  • Ensure platform stability, performance, security, and cost efficiency through proactive monitoring, tuning, retention management, and incident response
  • Partner with security, compliance, and architecture to enforce governance, access controls, and data policies
  • Create and maintain customer facing and service level documentation, and deliver training for the wider IT organization
  • Participate in an on‑call rotation and contribute to incident response and post incident reviews
Qualifications

To land this role you will need:

  • A track record of being self‑driven and solving complex problems with minimal oversight
  • 5+ years of experience in observability, SRE, platform, or infrastructure roles
  • Hands on experience with Datadog or a similar platform, including metrics, logs, APM and tracing, RUM, synthetics, dashboards, alerting, and service catalogs
  • Experience with open source observability tools such as Prometheus, Grafana, and InfluxDB, and strong hands on experience operating Elasticsearch or OpenSearch clusters in production, ideally on Kubernetes
  • Experience with index design, query tuning, and index lifecycle policies in Elasticsearch or OpenSearch, including troubleshooting performance and reliability issues
  • Practical knowledge of OpenTelemetry concepts and instrumentation patterns
  • Experience with ETL or data pipeline tools (for example Cribl Stream, Fluent Bit, Logstash, Kafka, or OpenTelemetry Collector)
  • Understanding of data schemas, normalization, and enrichment concepts for observability data
  • Familiarity with log routing, sampling, and cost optimization strategies
  • Good working experience with Terraform or similar IaC tooling, including using modules in team environments
  • Experience with Docker, Kubernetes, Git, and Helm
  • Experience with automation frameworks like Ansible and with integrating observability into CI and CD pipelines
  • Proficiency in a programming or scripting language such as Python or Go for APIs, automation, and integrations
  • Solid understanding of Linux fundamentals, networking basics, and security best practices
  • Strong collaboration and communication skills in a complex platform environment
  • A focus on reliability, automation, and measurable service outcomes
What makes you stand out:
  • You champion SLOs and build a culture of telemetry first engineering
  • You reduce alert fatigue with thoughtful design and continuous tuning
  • Experience operating observability at scale for high volume, low latency systems
  • History of leading migrations or modernization efforts, for example consolidating tools into Datadog or rolling out OpenTelemetry
  • Experience building self‑service observability enablement for developers and SREs
  • Awareness of hybrid environments and the implications for data residency, security, and cost
  • Demonstrated ability to balance retention, performance, and cost with clear analytics and guardrails
Education / Certification Requirements:
  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience
  • Preferred certifications: Datadog, Terraform Associate, Kubernetes certifications such as CKA or CKAD
  • Commitment to continual professional and technical development
Work environment:
  • FTE type of contract
  • Office location in São Paulo - Rua Joaquim Floriano - Rua Joaquim Floriano 413 SAO PAULO, São Paulo 04534-011 Brazil
  • Office location in Bogota - Avenida Carrera 9A - Avenida Carrera 9A # 115-06/30 Edificio Torre Tierrafirme Bogotá, 110111 Colombia
  • Hybrid model (4 days/week in the office, 1 day/week remote)
Benefits:
  • Medical and life insurance
  • Public Transportation Support
  • Meal and food allowances
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Monitoring and Observability Analyst Sr. (M-F)
Senior Monitoring and Observability Analyst Sr. (M-F)

Coderio • Bogotá

Presencial
COP 154.517.000 - 270.406.000
100% remote work
Long-term commitment
Collaborative international team
+1
Senior Sales Engineer - Key Accounts (AMER - East)
Senior Sales Engineer - Key Accounts (AMER - East)

Datadog • Floridablanca

Híbrido
COP 462.762.000 - 614.945.000
Onboarding
Global benefits
Mentor program
+3
Platform Observability Engineer — Hybrid (SP/Bogotá)
Platform Observability Engineer — Hybrid (SP/Bogotá)

StoneX Group Inc. • Bogotá

Híbrido
COP 310.578.000 - 465.867.000
Medical and life insurance
Public Transportation Support
Meal and food allowances
SRE - Observability Engineer
SRE - Observability Engineer

T-mapp Jobs • Bogotá

Presencial
COP 180.000.000 - 240.000.000
Competitive salary
Comprehensive health benefits
Continuous learning & certifications
+1
Software Engineer (Observability & Infrastructure)
Software Engineer (Observability & Infrastructure)

GFT Technologies • Colombia

A distancia
COP 147.994.000 - 221.993.000
Prepaid Medical Insurance
Rewards Program
Learning Academy
+1
Platform Operations Engineer (SRE) Colombia / Mexico
Platform Operations Engineer (SRE) Colombia / Mexico

Forte Group • Colombia

Híbrido
COP 226.714.000 - 302.287.000
Site Reliability Engineer
Site Reliability Engineer

DCT • Bogotá

A distancia
COP 156.225.000 - 234.339.000
Career Growth & Mentorship
Flexible Work Environment
Generative & Collaborative Culture
Ingeniero de Performance
Ingeniero de Performance

BigCheese • Colombia

A distancia
COP 96.741.000 - 135.439.000
20 días hábiles de vacaciones pagados
Clases de inglés
Regalos corporativos
+3
Senior/Specialist SRE, Colombia
Senior/Specialist SRE, Colombia

CI&T • Colombia

Presencial
COP 72.000.000 - 96.000.000
Maternity and Parental leaves
Mobile services subsidy
Sick pay – Life insurance
+3
Lead Operational Intelligence Engineer
Lead Operational Intelligence Engineer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 210.000.000
Learning culture
Health coverage
Medical leave coverage
+2