Senior Site Reliability Engineer

Teladoc Health

Buenos Aires

Presencial

ARS 3.200.000 - 5.200.000

Jornada completa

Hace 5 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Teladoc Health in Buenos Aires, Argentina is seeking a Senior Site Reliability Engineer with strong Azure experience to operate, scale, and improve the reliability of our cloud services. This role anchors an SRE team, partners with software engineering and operations, and mentors the team to embed reliability in the delivery lifecycle.

You will define SLIs/SLOs, improve observability, automate recovery, and ensure secure, cost-aware Azure operations across production workloads.

Responsabilidades

  • Define, implement, and improve Service Level Indicators, Service Level Objectives, and error budgets for critical applications and platform services.
  • Partner with application teams to improve service reliability, fault tolerance, scalability, and operational readiness.
  • Identify and eliminate recurring reliability issues through root cause analysis, automation, and architectural improvements.
  • Help design systems that are resilient to Azure region, zone, network, dependency, and deployment failures.
  • Participate in production readiness reviews for new services, major releases, and infrastructure changes.
  • Build and improve observability across applications, infrastructure, networks, and cloud services. Implement monitoring for the four golden signals of Latency, Traffic, Errors, and Saturation
  • Develop dashboards, alerts, logs, traces, and metrics using tools such as Azure Monitor, Log Analytics, Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic, or similar APM platforms
  • Create service health dashboards for engineering, operations, and leadership audiences.
  • Analyze system performance, bottlenecks, saturation trends, and capacity risks.
  • Improve backup, disaster recovery, failover, and business continuity practices.
  • Partner with engineering teams to implement resiliency patterns such as retries, circuit breakers, bulkheads, graceful degradation, and queue-based decoupling.
  • Support and improve production workloads running on Microsoft Azure.
  • Collaborate with cloud and network teams on secure, scalable Azure architecture.
  • Help enforce Azure operational standards, including tagging, monitoring, backup, recovery, identity, security, and cost awareness.
  • Conduct blameless post-incident reviews and document root causes, contributing factors, corrective actions, and prevention plans.
  • Work with Incident Management, NOC, Help Desk, and application teams to improve response processes and runbooks.
  • Work with Security Engineering to ensure production systems follow cloud security and compliance standards.
  • Support operational controls for identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.

Descripción del empleo

Join the team leading the next evolution of virtual care.

At Teladoc Health, you are empowered to bring your true self to work while helping millions of people live their healthiest lives.

Here you will be part of a high-performance culture where colleagues embrace challenges, drive transformative solutions, and create opportunities for growth. Together, we’re transforming how better health happens.

Position Summary

We are seeking a Senior Site Reliability Engineer with strong Microsoft Azure experience to help operate, scale, and improve the reliability of our critical cloud-based services. This role will anchor an SRE Team that will focus on production reliability, observability, automation, incident response, cloud operations, and continuous improvement.

The Sr. SRE partners with software engineering, product, security, and operations leadership to embed reliability principles into the software delivery lifecycle, mentors the SRE team, and serves as the technical authority for observability and reliability across mission‑critical healthcare workloads.

Role And Responsibilities

Reliability Engineering

  • Define, implement, and improve Service Level Indicators, Service Level Objectives, and error budgets for critical applications and platform services.
  • Partner with application teams to improve service reliability, fault tolerance, scalability, and operational readiness.
  • Identify and eliminate recurring reliability issues through root cause analysis, automation, and architectural improvements.
  • Help design systems that are resilient to Azure region, zone, network, dependency, and deployment failures.
  • Participate in production readiness reviews for new services, major releases, and infrastructure changes.

Observability and Monitoring

  • Build and improve observability across applications, infrastructure, networks, and cloud services. Implement monitoring for the four golden signals of Latency, Traffic, Errors, and Saturation
  • Develop dashboards, alerts, logs, traces, and metrics using tools such as Azure Monitor, Log Analytics, Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic, or similar APM platforms
  • Create service health dashboards for engineering, operations, and leadership audiences.

Performance, Capacity, and Resilience

  • Analyze system performance, bottlenecks, saturation trends, and capacity risks.
  • Improve backup, disaster recovery, failover, and business continuity practices.
  • Partner with engineering teams to implement resiliency patterns such as retries, circuit breakers, bulkheads, graceful degradation, and queue-based decoupling.

Azure Cloud Operations

  • Support and improve production workloads running on Microsoft Azure.
  • Collaborate with cloud and network teams on secure, scalable Azure architecture.
  • Help enforce Azure operational standards, including tagging, monitoring, backup, recovery, identity, security, and cost awareness.

Incident Management and Response

  • Conduct blameless post-incident reviews and document root causes, contributing factors, corrective actions, and prevention plans.
  • Work with Incident Management, NOC, Help Desk, and application teams to improve response processes and runbooks.

Security and Compliance Support

  • Work with Security Engineering to ensure production systems follow cloud security and compliance standards.
  • Support operational controls for identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.

As part of our hiring process, we verify identity and credentials, conduct interviews (live or video), and screen for fraud or misrepresentation. Applicants who falsify information will be disqualified.

Required Qualifications
  • 7+ years in site reliability, including hands‑on ownership of mission‑critical services, through a combination of applicable work experience, training, military experience, or education.
  • Deep Azure experience including Azure Monitor, Application Insights, AKS, and cloud‑native operations across hybrid infrastructure.
  • Proven track record designing and rolling out an SLO program with SLIs, SLOs, and error budget policy in production environments.
  • Observability Tools: Hands‑on experience with enterprise observability platform such as Datadog and Dynatrace, Elastic, Grafana, Prometheus or LogicMonitor.
  • Hands‑on experience implementing and configuring Datadog for monitoring, observability, and alerting.
Preferred Qualifications
  • Prior experience standing up or anchoring an SRE practice, including operating model, rituals, and adoption across multiple engineering teams.
  • Strong incident command experience and a track record of running blameless postmortems that drive measurable improvement.
  • Healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent compliance frameworks.
  • Multi‑cloud reliability experience (AWS) in addition to Azure.
  • Chaos engineering and resiliency testing experience (e.g., Gremlin, Chaos Mesh, Azure Chaos Studio).
  • Infrastructure as Code expertise (Terraform, Bicep, Ansible) for observability and remediation automation.
  • Scripting and programming proficiency (Python, PowerShell, Go) for automation, tooling, and integration work.
  • Experience with security controls, vulnerability management, compliance audits, and cloud governance.
  • Recognized industry certifications (e.g., Azure Solutions Architect, Google SRE certificate, CKA).
Why join Teladoc Health?
  • Teladoc Health is transforming how better health happens. Learn how when you join us in pursuit of our impactful mission.
  • Chart your career path with meaningful opportunities that empower you to grow, lead, and make a difference.
  • Join a multi-faceted community that celebrates each colleague’s unique perspective and is focused on continually improving, each and every day.
  • Contribute to an innovative culture where fresh ideas are valued as we increase access to care in new ways.
  • Enjoy an inclusive benefits program centered around you and your family, with tailored programs that address your unique needs.
  • Explore candidate resources with tips and tricks from Teladoc Health recruiters and learn more about our company culture by exploring #TeamTeladocHealth on LinkedIn.

As an Equal Opportunity Employer, we never have and never will discriminate against any job candidate or employee due to age, race, religion, color, ethnicity, national origin, gender, gender identity/expression, sexual orientation, membership in an employee organization, medical condition, family history, genetic information, veteran status, marital status, parental status, or pregnancy). In our innovative and inclusive workplace, we prohibit discrimination and harassment of any kind.

Teladoc Health respects your privacy and is committed to maintaining the confidentiality and security of your personal information. In furtherance of your employment relationship with Teladoc Health, we collect personal information responsibly and in accordance with applicable data privacy laws, including but not limited to, the California Consumer Privacy Act (CCPA). Personal information is defined as: Any information or set of information relating to you, including (a) all information that identifies you or could reasonably be used to identify you, and (b) all information that any applicable law treats as personal information. Teladoc Health’s Notice of Privacy Practices for U.S. Employees’ Personal information is available at this link.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Site Reliability Engineer - Argentina
Site Reliability Engineer - Argentina

Teladoc Health, Inc. • Argentina

Presencial
ARS 89.450.000 - 149.085.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

teladoc • Argentina

Presencial
ARS 135.663.000 - 195.957.000
Program Manager, Engineering Enablement - Argentina
Program Manager, Engineering Enablement - Argentina

Teladoc Health • Buenos Aires

Presencial
ARS 105.913.000 - 151.304.000
Senior Azure SRE: Reliability, Observability & Cloud Ops
Senior Azure SRE: Reliability, Observability & Cloud Ops

Smartek S.R.L • Argentina

A distancia
ARS 181.395.000 - 272.092.000
Director, Platform Engineering
Director, Platform Engineering

teladoc • Argentina

Híbrido
ARS 3.500.000 - 6.000.000
Azure SRE Lead: Reliability, Observability & Cloud Ops
Azure SRE Lead: Reliability, Observability & Cloud Ops

teladoc • Argentina

Presencial
ARS 135.663.000 - 195.957.000
Senior Azure SRE: Reliability, Observability & Scale
Senior Azure SRE: Reliability, Observability & Scale

Teladoc Health, Inc. • Argentina

A distancia
ARS 166.279.000 - 226.744.000
Azure SRE: Observability, Incident Response & Reliability
Azure SRE: Observability, Incident Response & Reliability

Teladoc Health, Inc. • Argentina

Presencial
ARS 89.450.000 - 149.085.000
DevOps Engineer III
DevOps Engineer III

Teladoc Health • Argentina

Presencial
ARS 2.000.000 - 2.800.000
Restaurant d'entreprise
Senior Azure SRE: Reliability, Observability & Scale
Senior Azure SRE: Reliability, Observability & Scale

Teladoc Health • Buenos Aires

Presencial
ARS 3.200.000 - 5.200.000