SRE (Engineering & Administration Background)

Fulcrum Digital

Ciudad de México

Híbrido

MXN 900.000 - 1.500.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Fulcrum Digital seeks a Site Reliability Engineer (SRE) in Mexico City on a hybrid schedule to own and optimize the health of our production environment. You will be at the center of platform reliability, collaborating with a global team across multiple time zones.

You will plan, monitor, and improve production systems, respond to incidents, automate deployments, and drive DevOps practices across CI/CD pipelines while aligning with ITIL/ITSM processes and capacity planning.

Formación

  • Significant monitoring experience with SOP creation; Splunk and Dynatrace required.
  • Strong hands-on Linux experience and shell scripting proficiency.
  • Solid ITIL/ITSM understanding and incident response skills.
  • Experience supporting CI/CD pipelines and DevOps best practices.
  • Ability to analyze, troubleshoot, and drive improvements across environments.

Responsabilidades

  • Plan, manage, and oversee all aspects of a production environment.
  • Define strategies for application performance monitoring and optimization.
  • Respond to incidents, drive platform improvements, and measure MTTR.
  • Support code deployment across multiple environments with automation focus.
  • Design, develop, and standardize monitoring and alerting.
  • Own full service lifecycle from design to deployment and refinement.
  • Analyze ITSM activity and provide feedback to development teams.
  • Support pre-launch activities including capacity planning and reviews.
  • Support CI/CD pipelines through validation and gating.
  • Monitor availability and latency to keep services running smoothly.
  • Drive scalable, reliability-focused system changes.
  • Perform root cause analysis and provide on-call support on rotation.
  • Collaborate with global teams across multiple time zones.
  • Share knowledge and mentor others on processes and procedures.
  • Occasional off-hours work required.

Conocimientos

Monitoring
Linux
Shell scripting
ITIL/ITSM
Incident response
DevOps
CI/CD
Capacity planning
Troubleshooting
Cross-team collaboration

Herramientas

Splunk
Dynatrace
Prometheus
Grafana
Git
Bitbucket

Descripción del empleo

Fulcrum Digital is a global AI-first enterprise transformation company with over 25 years of experience. We partner with enterprises across financial services, insurance, healthcare, retail, manufacturing, higher education, and logistics to move from AI experimentation to scalable business outcomes. With over 100 global clients, including Fortune 500 enterprises, we combine deep industry expertise with capabilities in enterprise AI, digital engineering, cloud modernisation, platform integration, and generative AI.

The Role

We are looking for a Site Reliability Engineer (SRE), based in Mexico City on a hybrid schedule, to own and optimize the health of our production environment. You will be at the center of platform reliability, working across the full service lifecycle from design through operation, while collaborating with a global team across multiple time zones.

What You'll Do

  • Plan, manage, and oversee all aspects of a production environment
  • Define strategies for application performance monitoring and optimization in production
  • Respond to incidents, drive platform improvements, and measure incident reduction over time
  • Support code deployment across multiple lower environments, with a strong focus on automation
  • Design, develop, and standardize monitoring and alerting mechanisms
  • Take a holistic, cross-stack approach to problem-solving during production events to optimize mean time to recovery (MTTR)
  • Own the full service lifecycle, from inception and design through deployment, operation, and refinement
  • Analyze ITSM activity and provide feedback loops to development teams on operational gaps and resiliency concerns
  • Support pre-launch activities, including system design consulting, capacity planning, and launch reviews
  • Support CI/CD pipelines through validation and operational gating, championing DevOps best practices
  • Monitor availability, latency, and overall system health to keep services running smoothly
  • Drive sustainable scaling through automation and reliability-focused system changes
  • Perform root cause analysis and on-call support on a rotational basis
  • Collaborate with a global team spread across multiple tech hubs and time zones
  • Share knowledge and mentor others on processes and procedures
  • Occasional off-hours work required
Requirements

Must Have

  • Significant strength in monitoring, including SOP creation — Splunk and Dynatrace experience is required
  • Strong hands-on experience with Linux
  • Shell scripting proficiency
  • Strong application troubleshooting skills
  • Solid understanding of ITIL / ITSM processes

Also Required

  • SQL knowledge
  • Working knowledge of Git / Bitbucket
  • Root cause analysis and capacity planning experience

Preferred Qualifications

  • Card payment knowledge (payment flows, switching, settlements, authorization flows)
  • Experience with Prometheus/Grafana
  • Cloud experience — AWS and/or Azure
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Hybrid SRE: Production Reliability & Automation
Hybrid SRE: Production Reliability & Automation

Fulcrum Digital • Ciudad de México

Híbrido
MXN 900.000 - 1.500.000
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

Híbrido
MXN 1.049.000 - 1.400.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

Híbrido
MXN 870.000 - 1.306.000
Professional growth
Competitive compensation
Exciting projects
+1
SRE | On site in GDL or CDMX
SRE | On site in GDL or CDMX

gsbsolutions1 • Región Centro

Presencial
MXN 72.000 - 88.000
Excellent superior benefits
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

A distancia
MXN 1.433.000 - 1.793.000
Sr. Manager SRE (Individual Contributor)
Sr. Manager SRE (Individual Contributor)

Capital One • Ciudad de México

Híbrido
MXN 1.400.000 - 2.100.000
Senior AWS Site Reliability Engineers - 2850
Senior AWS Site Reliability Engineers - 2850

Xideral • Región Centro

Híbrido
MXN 1.200.000 - 1.500.000
Premium Benefits
Performance bonuses
SGMM Medical insurance
SRE: Cloud, CI/CD & Automation (On-site)
SRE: Cloud, CI/CD & Automation (On-site)

gsbsolutions1 • Región Centro

Presencial
MXN 72.000 - 88.000
Excellent superior benefits
Customer Service Reliability Engineer
Customer Service Reliability Engineer

Thales • Ciudad de México

Híbrido
MXN 900.000 - 1.500.000
Senior SRE Engineer: Reliability, Automation & Growth
Senior SRE Engineer: Reliability, Automation & Growth

Spin Careers • Ciudad de México

Presencial
MXN 900.000 - 1.300.000