Sre +Dynatrace | Mexico

The Photon Group

Región Centro

Presencial

MXN 900.000 - 1.400.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto — currículum adaptado y carta de presentación, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

The Photon Group está buscando un SRE senior para Guadalajara, México, orientado a fiabilidad de sistemas distribuidos y servicios a gran escala. Trabajarás en diseño de SLI/SLO, observabilidad y mejora continua de operaciones en entornos Cloud.

Se valorará experiencia en Dynatrace, Prometheus, Grafana y herramientas de monitoreo, así como capacidad de liderar incidentes y colaborar con equipos de producto e ingeniería. Proyecto desafiante y oportunidad de crecimiento.

Formación

  • Se requiere experiencia en SRE/ingeniería de software o operaciones de producción para plataformas a gran escala.
  • Capacidad para definir y medir SLIs/SLOs y modelos de fiabilidad basados en SLOs.
  • Conocimiento de plataformas en la nube (AWS, Azure o Google Cloud) y herramientas de observabilidad.

Responsabilidades

  • Liderar incidentes y servir como punto de escalada para incidentes de alto impacto.
  • Diseñar, implementar y mantener plataformas de observabilidad (métricas, logs, trazas, telemetría).
  • Trabajar con equipos para reducir toil operativo y convertir trabajo repetitivo en historias/épicas en Jira.
  • Desarrollar resiliencia mediante degradación controlada, conmutación por error y recuperación ante desastres.

Conocimientos

SRE/DevOps
Distributed systems
Java/J2EE
React (bonus)
SLO/SLA design
Incident response
CI/CD pipelines
Jira backlog
Networking basics
Linux
Cloud platforms

Educación

Bachelor's Degree

Herramientas

Dynatrace
Prometheus
Grafana
ELK
Akamai
New Relic
Datadog

Descripción del empleo

SRE +Dynatrace | Mexico

Empresa : The Photon Group Tipo de empleo : Tiempo completo Mexico

Descripción del trabajo - SRE +Dynatrace | Mexico
Description

Location: Guadalajara (Mexico)

What you'll do:

  • Experience of working with large scale distributed systems, including scalability, disaster recovery and fault tolerance.
  • Expertise Python scripting .
  • Define, implement, and own SLIs, SLOs, and error budgets for critical microservices in collaboration with product and engineering teams.
  • Use error budgets to influence release decisions, prioritize reliability work, and manage operational risk.
  • Design and maintain observability platforms including metrics, logs, traces, and real-time telemetry.
  • Track, manage, and reduce operational toil by converting repetitive operational work into Jira stories and epics with clear ownership and measurable outcomes.
  • Design, implement, and validate resiliency mechanisms such as graceful degradation, redundancy, automated failover, and disaster recovery.
  • Lead incident response, act as an escalation point for high-severity incidents, and drive blameless postmortems.
  • Partner with scrum teams to improve reliability through release readiness reviews, production change validation, and testing strategies.
  • Capture incident action items and reliability improvements in Jira, ensuring closure, accountability, and continuous improvement.
  • Perform deep root cause analysis, debugging, and performance tuning across distributed systems.
  • Provide technical leadership and mentoring to junior SREs and engineers.
  • Promote shift-left reliability by embedding operability, monitoring, and failure testing early in the SDLC.
  • Strong knowledge on CICD Pipeline, GIT, AWS/Azure/GCP as Paas service
  • Demonstrated knowledge of Configuration Management and Deployment tools automation
  • Strong Experience with networking concepts and protocols (HTTP, HTTPS, Telnet, SSH, Firewall, VPN, Routing and Load Balancing)
  • Strong Experience with Linux
  • Experience with Monitoring solutions like Prometheus, Grafana, Products like ELK/Splunk etc.
  • Experience of working with large scale systems
  • Experience with containers and orchestration technologies like Docker, Kubernetes
  • Experience on Service Mesh like Istio, etc. would be added Advantage
  • Experience with any CDN like Akamai etc..

What you'll bring:

  • Bachelor's Degree in Computer Science or related technical field.
  • 4+ years of experience in SRE, software engineering, or production operations supporting large-scale eCommerce platforms.
  • Hands‑on experience with Java/J2EE-based distributed systems. React experience is a plus.
  • Proven ability to design and operate systems using SLO-driven reliability models.
  • Experience defining and measuring SLIs (availability, latency, error rates, throughput, saturation).
  • Good understanding with NoSQL technologies and RDBMS. Should be able to write queries to fetch results from database.
  • Experience deploying and operating services on cloud platforms (AWS, Azure, or Google Cloud).
  • Expertise with observability, APM, and caching tools (Dynatrace, Splunk, ELK, Akamai, QuantumMetric/Tealeaf, etc.).
  • Strong experience using Jira for backlog management, incident follow-ups, toil reduction tracking, and cross-team coordination.
  • Ability to independently own services and drive reliability initiatives end-to-end.
  • Strong communication skills and ability to influence engineering and product teams.
  • Experience being on On-Call rotation and handling critical/high incidents.

Good to have:

  • Candidates with application support experience can be considered.
  • Any monitoring tools experience is acceptable such as New Relic or Datadog can also be considered.
  • Candidates with 3 to 4 years of experience are fine; even junior resources with 3 years of experience can be considered.
  • Akamai experience is optional.
  • Any cloud experience is acceptable.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

SRE + Dynatrace: Reliability & Observability Leader
SRE + Dynatrace: Reliability & Observability Leader

The Photon Group • Región Centro

Presencial
MXN 900.000 - 1.400.000
SRE | On site in GDL or CDMX
SRE | On site in GDL or CDMX

gsbsolutions1 • Región Centro

Presencial
MXN 72.000 - 88.000
Excellent superior benefits
SRE (Engineering & Administration Background)
SRE (Engineering & Administration Background)

Fulcrum Digital • Ciudad de México

Híbrido
MXN 900.000 - 1.500.000
Lead SRE Engineer
Lead SRE Engineer

Cloudsufi • Región Centro

Presencial
MXN 900.000 - 1.400.000
Dynatrace Senior Engineer
Dynatrace Senior Engineer

Sequoia Connect LLC • Xico

Híbrido
MXN 900.000 - 1.200.000
Hybrid work model
Dynatrace Engineer
Dynatrace Engineer

Sequoia Connect LLC • Saltillo

Híbrido
MXN 420.000 - 640.000
Site Reliability Engineering (SRE) Lead - 2770
Site Reliability Engineering (SRE) Lead - 2770

Xideral • Región Centro

A distancia
MXN 700.000 - 900.000
Attractive Salary
Performance bonuses
SGMM Medical insurance
App Support Engineer Web-Based Pr
App Support Engineer Web-Based Pr

Softtek • Ciudad de México

Presencial
MXN 400.000 - 540.000
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

Presencial
MXN 1.433.948 - 1.792.436
Senior Site Reliability Engineer
Senior Site Reliability Engineer

adlytics GmbH • Región Centro

Híbrido
MXN 900.000 - 1.300.000
Sueldo acorde a experiencia
Esquema 100% nómina
Prestaciones de Ley
+5