Senior Site Reliability Engineer (Azure)

OpsGenius

Colombia

Presencial

COP 380.409.000 - 475.511.000

Jornada completa

Hace 7 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Paid time off
Health insurance
Learning credits
Performance incentives
Career support
Fully remote

Descripción de la vacante

OpsGenius is hiring a Senior Site Reliability Engineer to build the observability, on-call, and deployment foundation for a large Azure-based SaaS platform. You will own critical operational areas, from performance tuning to cost optimization, in a remote LATAM-friendly setup.

You will collaborate with a US-based team and drive production reliability through IaC, automated testing, and robust runbooks, ensuring high uptime and scalable cloud infrastructure.

Formación

  • 7+ years in SRE, DevOps, or cloud infrastructure with on-call production ownership.
  • Strong Azure experience; AWS background with light Azure exposure is not a fit.
  • Hands-on Azure SQL DBA experience: indexing, query tuning, HA/DR, backup/restore.
  • Comfort reading query execution plans and diagnosing DB performance.
  • Terraform or equivalent IaC in production.
  • Experience with observability tooling: Application Insights, Azure Monitor, Datadog.
  • Strong written and spoken English; daily communication with US team.
  • Work schedule overlapping US Eastern hours.

Responsabilidades

  • Build out observability: SLOs, alerts, dashboards using Azure Monitor and App Insights.
  • Participate in on-call rotation, paging tooling, escalation paths, incident docs.
  • Modernize deployments with blue/green/canary, automated tests, Terraform IaC.
  • Tune Azure SQL performance across multi-tenant estate.
  • Own backup/restore strategy for Azure SQL; test PITR and drills.
  • Optimize Cosmos DB RU consumption and partition keys for throughput.
  • Operate containerized workloads on Azure Container Apps and App Service.
  • Write runbooks for incident triage, rollback, restore, secret rotation.

Conocimientos

Azure expertise
On-call ownership
Observability tooling
English communication

Herramientas

Terraform
Azure Monitor
Application Insights
Datadog
Cosmos DB
Azure SQL

Descripción del empleo

Senior Site Reliability Engineer (Azure)

Location: Remote, LATAM

Job Type: Full-Time, with benefits

OpsGenius is a boutique US-based cloud operations firm. We take over the operational side of large production platforms so engineering teams can get back to building.

We are hiring a Senior SRE to join a small, senior team standing up the operational foundation for a large multi-tenant SaaS platform on Azure, running at significant scale across many environments and a high-volume API surface.

The platform has scaled quickly, and you will be one of the people building the operational foundation alongside it: observability, on-call, and deployment practices. This is hands-on production ownership, not a ticket queue.

WHAT YOU WILL DO
  • Build out observability: SLO definitions, alert rationalization, synthetic checks, and dashboards using Azure Monitor, Application Insights, and potentially Datadog
  • Stand up and participate in an on-call rotation, including paging tooling, escalation paths, and incident documentation
  • Modernize deployments: blue/green or canary patterns with validated rollback, automated smoke tests, and Infrastructure as Code with Terraform
  • Perform Azure SQL performance work: query optimization, indexing strategy, and execution plan analysis across a large multi-tenant estate
  • Own backup and restore strategy for Azure SQL, including point-in-time recovery testing and periodic restore drills
  • Tune elastic pool sizing and evaluate DTU versus vCore tradeoffs for cost and performance
  • Optimize Cosmos DB RU consumption and partition key design for high-throughput workloads
  • Operate containerized workloads on Azure Container Apps, alongside App Service and Service Bus
  • Write operational runbooks (incident triage, rollback, backup restore, secret rotation, certificate renewal) that any on-call engineer can follow
  • Contribute to Azure cost optimization: right-sizing, autoscaling tuning, tagging, and cost reporting
  • Support security and reliability hardening over time, including IAM reviews, backup restore drills, and DR exercises
  • Collaborate daily with a US-based team during overlapping hours
WHAT WE ARE LOOKING FOR

Required

  • 7+ years in SRE, DevOps, or cloud infrastructure roles, including direct ownership of production systems under an on-call rotation
  • Deep, hands-on Azure experience. AWS-primary backgrounds with light Azure exposure will not be a fit
  • Hands-on Azure SQL DBA experience: indexing, query tuning, HA/DR (failover groups, geo-replication), and backup/restore. This is not a generalist SRE role; real database ownership is expected
  • Comfort reading query execution plans and diagnosing performance regressions at the database level, not just infrastructure monitoring
  • Terraform or equivalent IaC in production
  • Experience with observability tooling: Application Insights, Azure Monitor, Datadog, or similar
  • Strong written and spoken English; daily communication with a US team and occasional client stakeholders
  • Work schedule overlapping US Eastern hours (roughly 9 to 5 Eastern is ideal)
Nice to Have
  • Experience in HIPAA, PHI, or other regulated environments. A background check to healthcare-industry standards is required prior to production access
  • Cosmos DB RU optimization and partition key design at high throughput
  • Prior work in a multi-tenant SaaS environment
ON-CALL EXPECTATIONS

This role includes a pager-based on-call rotation covering SEV-1 and SEV-2 incidents, shared with the rest of the SRE team. On-call is a core part of the role. Expect it to be light in the first month and ramp as the team takes over production responsibility.

COMPLIANCE AND ACCESS
  • All personnel are named and approved by the client before any access is provisioned
  • Background checks to healthcare-industry standards are completed prior to production access
  • Production access is provisioned through the client identity provider with MFA and time-bound elevation
BENEFITS
  • Paid time off
  • Supplemental health insurance
  • Learning credits for training and certification
  • Performance incentives
  • Regular one on ones and ongoing career support
  • Fully remote
WHY THIS ROLE
  • You are building the observability, on-call, and deployment practices, not inheriting someone else’s
  • Small senior team, direct access to senior engineers, no layers of process
  • Interesting scale problems: multi-tenant data at volume and real performance challenges, in an environment that has to stay up while you harden it
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Azure SRE - Observability & High-Scale SaaS
Senior Azure SRE - Observability & High-Scale SaaS

OpsGenius • Colombia

Presencial
COP 380.409.000 - 475.511.000
Paid time off
Health insurance
Learning credits
+3
Site Reliability Engineer
Site Reliability Engineer

AgileEngine • Colombia

Presencial
COP 283.921.000 - 441.654.000
Professional growth
Competitive USD compensation
Flextime
+1
Site Reliability Engineer
Site Reliability Engineer

Source Meridian • Colombia

Presencial
COP 244.021.000 - 406.703.000
Workout bonus
Home office setup bonus
Learning platform bonus
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Home office
Career plans to allow for extensive成长
International projects
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Colombia

Presencial
COP 284.270.000 - 473.784.000
Site Reliability Engineer (Cloud)
Site Reliability Engineer (Cloud)

Sharesource Australia BPO Corporation • Norte

Híbrido
COP 110.000.000 - 180.000.000
Remote + Hybrid
Work-life balance
Open culture
+3
Senior Multi-Cloud SRE & Incident Commander
Senior Multi-Cloud SRE & Incident Commander

AgileEngine • Colombia

Presencial
COP 283.921.000 - 441.654.000
Professional growth
Competitive USD compensation
Flextime
+1
Site Reliability Engineer
Site Reliability Engineer

Sharesource • Norte

Híbrido
COP 110.000.000 - 170.000.000
Remote + Hybrid Flexibility
Work‑Life Balance
Open Culture
+2