Site Reliability Engineer III

Teladoc Health

Argentina

Híbrido

ARS 1.200.000 - 2.200.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Teladoc Health in Argentina is seeking a Site Reliability Engineer (SRE) with deep Azure expertise, focused on Observability, Monitoring, and Incident Response to ensure availability and performance of hybrid cloud services.

You will design observability frameworks, automate monitoring/alerting, build runbooks, and lead post-incident retrospectives across multi-cloud environments, collaborating with engineering, security, and operations teams.

Formación

  • 3+ years of SRE or related experience.
  • Hands-on experience with enterprise observability platforms.
  • Strong knowledge of cloud-native operations.

Responsabilidades

  • Design, implement, and maintain observability across Azure and multi-cloud environments.
  • Define and standardize SLIs/SLOs/SLAs to measure service health and user experience.
  • Develop dashboards and automated alerting to detect degradations.
  • Create and maintain on-call runbooks and playbooks to reduce MTTR.
  • Lead post-incident blameless retrospectives and continuous improvement.
  • Develop automation for self-healing systems and remediation workflows.
  • Contribute to disaster recovery and business continuity planning across multi-clouds.
  • Collaborate with security, network, and system teams to ensure observability is aligned with governance.
  • Mentor staff in observability tools, monitoring strategies, and incident management.

Conocimientos

SRE fundamentals
Incident management
Automation
Collaboration
Communication

Herramientas

Datadog
Dynatrace
Grafana
Prometheus
Elastic
Azure Monitor

Descripción del empleo

We are seeking a highly skilled Site Reliability Engineer (SRE) with deep experience in Azure environments, specializing in Observability, Monitoring, and Incident Response. This role is critical to ensuring the availability, reliability, and performance of our hybrid cloud infrastructure and services. The ideal candidate will design and implement observability frameworks, drive automation in monitoring and alerting, and lead effective incident management processes across multi-cloud environments.

This position requires strong technical acumen in cloud-native operations, a proactive mindset toward reliability engineering, and the ability to collaborate with engineering, operations, and security teams to maintain mission-critical healthcare and enterprise workloads.

Essential Duties and Responsibilities
Observability & Monitoring
  • Design, implement, and maintain observability solutions across Azure (e.g., Azure Monitor, Datadog, Grafana-Prometheus, Dynatrace, Elastic).
  • Define and standardize SLIs/SLOs/SLAs to measure service health and customer experience.
  • Develop dashboards and automated alerting to proactively identify service degradations.
  • Build and maintain on-call runbooks and playbooks to reduce time-to-resolution.
  • Drive post-incident “blameless” retrospectives and continuous improvement initiatives.
Reliability Engineering
  • Develop automation for self-healing systems, monitoring remediation, and incident mitigation.
  • Contribute to disaster recovery and business continuity planning across multi-cloud platforms.
  • Work with engineering teams to design for resiliency, scalability, and reliability from the ground up.
  • Partner with security, network, and system engineering teams to ensure observability integrates with compliance and governance frameworks.
  • Advocate for best practices in cloud-native reliability engineering.
  • Mentor engineering staff in observability tools, monitoring strategies, and incident management.

The time spent on each responsibility reflects an estimate and is subject to change dependent on business needs.

Supervisory Responsibilities

No

Required Qualifications
  • +3 years of experience, or equivalent demonstrated through a combination of work experience, training, military experience, or education for:
  • Observability Tools: Hands-on experience with enterprise observability platform such as Datadog and Dynatrace, Elastic, Grafana, Prometheus orLogicMonitor.
  • Monitoring & Alerting: Deep understanding of metrics, logs, traces, and distributed system monitoring.
Preferred Qualifications
  • Automation & Infrastructure as Code (IaC): Proficiency with Terraform, Bicep, Ansible, or similar tools to automate monitoring and remediation workflows.
  • Kubernetes Observability: Knowledge of AKS logging, tracing, and monitoring in containerized environments.
  • AI & Observability: Exposure to AI-driven monitoring, anomaly detection, or predictive alerting tools.
  • Programming/Scripting: Scripting skills in Python, PowerShell, or similar languages for automation and tool integration.
  • Chaos Engineering: Experience with resiliency testing and tools such as Gremlin or Chaos Mesh.
  • Healthcare & Compliance: Experience in healthcare IT environments with HIPAA, HITRUST, or other compliance frameworks.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Teladoc Health • Argentina

Presencial
ARS 1.800.000 - 2.600.000
Azure SRE: Observability, Incident Response & Reliability
Azure SRE: Observability, Incident Response & Reliability

Teladoc Health, Inc. • Argentina

Presencial
ARS 89.450.000 - 149.085.000
Site Reliability Engineer - Argentina
Site Reliability Engineer - Argentina

teladoc • Argentina

Presencial
ARS 1.200.000 - 1.800.000
Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 89.314.000 - 119.087.000
Azure SRE: Observability, Incident Response & Reliability
Azure SRE: Observability, Incident Response & Reliability

teladoc • Argentina

Presencial
ARS 1.200.000 - 1.800.000
Site Reliability Engineer - Argentina
Site Reliability Engineer - Argentina

Teladoc Health, Inc. • Argentina

Presencial
ARS 89.450.000 - 149.085.000
Site Reliability Engineer
Site Reliability Engineer

EPAM Systems • Argentina

Presencial
ARS 134.672.073 - 179.562.764
Azure SRE: Observability & Incident Response Lead
Azure SRE: Observability & Incident Response Lead

Teladoc Health • Argentina

Híbrido
ARS 1.200.000 - 2.200.000
Senior Azure SRE: Reliability & Observability Lead
Senior Azure SRE: Reliability & Observability Lead

Teladoc Health • Argentina

Presencial
ARS 1.800.000 - 2.600.000
OT/PCN Site Reliability Engineer
OT/PCN Site Reliability Engineer

Chevron - Internal • Partido de Vicente López

Presencial
ARS 59.543.000 - 89.315.000