Senior Site Reliability Engineer

Teladoc Health

Argentina

Presencial

ARS 136.170.000 - 196.690.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Descripción de la vacante

Teladoc Health is seeking a Senior Site Reliability Engineer to anchor an SRE team focused on reliability, observability, automation, and cloud operations for healthcare workloads.

The role partners with software engineering, product, security, and operations leadership to embed reliability principles across the delivery lifecycle and mentors the team in observability best practices.

Formación

  • 7+ years in site reliability with ownership of mission-critical services.
  • Deep Azure experience including Monitor, Insights, AKS, and cloud-native operations.
  • Experience designing and rolling out an SLO program with SLIs and error budgets.

Responsabilidades

  • Define and improve service level objectives and error budgets for critical apps and platform services.
  • Partner with application teams to improve reliability, scalability, and readiness.
  • Lead root-cause analysis and automation to eliminate recurring issues.

Conocimientos

Azure
Observability
SRE practices

Descripción del empleo

We are seeking a Senior Site Reliability Engineer (Sr. SRE) with strong Microsoft Azure experience to help operate, scale, and improve the reliability of our critical cloud-based services. This role will anchor an SRE Team that will focus on production reliability, observability, automation, incident response, cloud operations, and continuous improvement.

The Sr. SRE partners with software engineering, product, security, and operations leadership to embed reliability principles into the software delivery lifecycle, mentors the SRE team, and serves as the technical authority for observability and reliability across mission‑critical healthcare workloads.

Role and Responsibilities
Reliability Engineering
  • Define, implement, and improve Service Level Indicators, Service Level Objectives, anderror budgetsfor critical applications and platform services.
  • Partner with application teams to improve service reliability, fault tolerance, scalability, and operational readiness.
  • Identify and eliminate recurring reliability issues through root cause analysis, automation, and architectural improvements.
  • Help design systems that are resilient to Azure region, zone, network, dependency, and deployment failures.
  • Participate in production readiness reviews for new services, major releases, and infrastructure changes.
Observability and Monitoring
  • Build and improve observability across applications, infrastructure, networks, and cloud services. Implement monitoring for the four golden signals of Latency, Traffic, Errors, and Saturation
  • Develop dashboards, alerts, logs, traces, and metrics using tools such as Azure Monitor, Log Analytics, Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic, or similar APM platforms
  • Create service health dashboards for engineering, operations, and leadership audiences.
Performance, Capacity, and Resilience
  • Analyze system performance, bottlenecks, saturation trends, and capacity risks.
  • Improve backup, disaster recovery, failover, and business continuity practices.
  • Partner with engineering teams to implement resiliency patterns such as retries, circuit breakers, bulkheads, graceful degradation, and queue‑based decoupling.
Azure Cloud Operations
  • Support and improve production workloads running on Microsoft Azure.
  • Collaborate with cloud and network teams on secure, scalable Azure architecture.
  • Help enforce Azure operational standards, including tagging, monitoring, backup, recovery, identity, security, and cost awareness.
Incident Management and Response
  • Conduct blameless post‑incident reviews and document root causes, contributing factors, corrective actions, and prevention plans.
  • Work with Incident Management, NOC, Help Desk, and application teams to improve response processes and runbooks.
Security and Compliance Support
  • Work with Security Engineering to ensure production systems follow cloud security and compliance standards.
  • Support operational controls for identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.
Required Qualifications
  • 7+ years in site reliability, including hands‑on ownership of mission‑critical services, through a combination of applicable work experience, training, military experience, or education.
  • Deep Azure experience including Azure Monitor, Application Insights, AKS, and cloud‑native operations across hybrid infrastructure.
  • Proven track record designing and rolling out an SLO program with SLIs, SLOs, and error budget policy in production environments.
Preferred Qualifications
  • Prior experience standing up or anchoring an SRE practice, including operating model, rituals, and adoption across multiple engineering teams.
  • Strong background in observability, including APM (Datadog, Dynatrace, New Relic), Prometheus, Grafana, and OpenTelemetry.
  • Strong incident command experience and a track record of running blameless postmortems that drive measurable improvement.
  • Healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent compliance frameworks.
  • Multi‑cloud reliability experience (AWS) in addition to Azure.
  • Chaos engineering and resiliency testing experience (e.g., Gremlin, Chaos Mesh, Azure Chaos Studio).
  • Infrastructure as Code expertise (Terraform, Bicep, Ansible) for observability and remediation automation.
  • Scripting and programming proficiency (Python, PowerShell, Go) for automation, tooling, and integration work.
  • Experience with security controls, vulnerability management, compliance audits, and cloud governance.
  • Recognized industry certifications (e.g., Azure Solutions Architect, Google SRE certificate, CKA).
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Azure SRE: Reliability, Observability & Scale
Senior Azure SRE: Reliability, Observability & Scale

Teladoc Health, Inc. • Argentina

A distancia
ARS 166.279.000 - 226.744.000
Senior Azure SRE: Reliability, Observability & Cloud Ops
Senior Azure SRE: Reliability, Observability & Cloud Ops

Smartek S.R.L • Argentina

A distancia
ARS 181.395.000 - 272.092.000
Site Reliability Engineer
Site Reliability Engineer

AgileEngine • Argentina

Híbrido
ARS 2.000.000 - 4.000.000
Professional growth
Competitive USD-based compensation
Flextime
+1
Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 3.500.000 - 6.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Teladoc Health • Buenos Aires

Presencial
ARS 3.200.000 - 5.200.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off
Upskilling programs
+2
Senior Azure SRE: Reliability, Observability & Scale
Senior Azure SRE: Reliability, Observability & Scale

Teladoc Health • Buenos Aires

Presencial
ARS 3.200.000 - 5.200.000
Senior SRE — Remote, Multi-Cloud & Incident Commander
Senior SRE — Remote, Multi-Cloud & Incident Commander

AgileEngine • Argentina

Híbrido
ARS 2.000.000 - 4.000.000
Professional growth
Competitive USD-based compensation
Flextime
+1
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Visa Hunt • Argentina

Híbrido
ARS 4.000.000 - 7.000.000
Flexible remote/office options
Competitive salary and compensation
Career growth & mentorship
+2
Senior SRE: Remote-Ready, Scalable Cloud Reliability
Senior SRE: Remote-Ready, Scalable Cloud Reliability

Visa Hunt • Argentina

Presencial
ARS 181.190.000 - 271.784.000
Flexible working format
Education reimbursement
Mentorship program
+2