SRE

Fulcrum Digital

Ciudad de México

Remote

MXN 600,000 - 1,000,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Fulcrum Digital is seeking a Site Reliability Engineer to own the health, stability, and performance of production environments for a client engagement. You will define monitoring strategies, drive automation across deployments and operations, and collaborate with development teams to reduce incidents and improve resiliency.

Join a global team across time zones, implement scalable observability solutions, and ensure smooth code deployments with a strong service ownership mindset and proactive

Qualifications

  • Experience with monitoring tools such as Splunk or Dynatrace.
  • Working knowledge of ITIL/ITSM practices.
  • Strong troubleshooting across complex, multi-layered platforms.
  • Proficiency in SQL and PL/SQL.
  • Experience with Jenkins and CI/CD pipelines.
  • Hands-on scripting with Groovy, YAML, and Shell.
  • Experience with Git and Bitbucket.
  • Hands-on experience with Kubernetes and AWS.
  • Proven experience in production support leadership, including runbook and support model creation.
  • Experience defining monitoring and alerting strategies.
  • Experience with disaster recovery and resiliency planning.
  • Experience with deployment readiness validation and operational process design.
  • Solid background in root cause analysis and problem management.

Responsibilities

  • Plan, manage, and oversee all aspects of the production environment.
  • Define strategies for application performance monitoring and optimization in production.
  • Design, develop, and standardize monitoring and alerting mechanisms for supported applications.
  • Respond to incidents, improve the platform based on feedback, and measure the reduction of incidents over time.
  • Take a holistic approach to problem solving during production events, connecting the dots across the full technology stack to optimize mean time to recover (MTTR).
  • Analyze ITSM activities for the platform and provide a feedback loop to development teams on operational gaps or resiliency concerns.
  • Engage in and improve the whole lifecycle of services, from inception and design through deployment, operation, and refinement.
  • Support services before they go live through system design consulting, capacity planning, and launch reviews.
  • Maintain live services by measuring and monitoring availability, latency, and overall system health.
  • Support code deployments into multiple lower environments, supporting current processes while automating wherever possible.
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead on DevOps automation and best practices.
  • Scale systems sustainably through automation, pushing for changes that improve both reliability and velocity.
  • Collaborate with a global team spread across tech hubs in multiple geographies and time zones, sharing knowledge and explaining processes and procedures to others.

Skills

Linux
ITIL/ITSM
Troubleshooting
SQL/PLSQL
CI/CD (Jenkins)
Scripting (Groovy, YAML, Shell)
Git/Bitbucket
Kubernetes
AWS
Production support leadership
Monitoring strategy
Disaster recovery
DevOps automation
Communication

Tools

Splunk
Dynatrace
Jenkins
Kubernetes
AWS
Git
Bitbucket

Job description

Fulcrum Digital is a globalAI-first enterprise transformation company with over 25 years of experience. Wepartner with enterprises across financial services, insurance, healthcare,retail, manufacturing, higher education, and logistics to move from AI experimentationto scalable business outcomes. Fulcrum Digital works with over 100 globalclients, including Fortune 500 enterprises, combining deep industry expertisewith capabilities in enterprise AI, digital engineering, cloud modernisation,platform integration, and generative AI.

The Role

As a Site Reliability Engineer,you will own the health, stability, and performance of a production environmentsupporting one of our client engagements. You will define how applications aremonitored and supported, drive automation across deployment and operations, andwork closely with development teams to reduce incidents and improve resiliencyover time. This role suits someone with a strong service ownership mindset whoenjoys connecting the dots across a complex technology stack and working with aglobal team across multiple time zones.

What You'll Do
  • Plan, manage, and oversee all aspects of the productionenvironment.
  • Define strategies for application performancemonitoring and optimization in production.
  • Design, develop, and standardize monitoring andalerting mechanisms for supported applications.
  • Respond to incidents, improve the platform based onfeedback, and measure the reduction of incidents over time.
  • Take a holistic approach to problem solving duringproduction events, connecting the dots across the full technology stack tooptimize mean time to recover (MTTR).
  • Analyze ITSM activities for the platform and provide afeedback loop to development teams on operational gaps or resiliency concerns.
  • Engage in and improve the whole lifecycle of services,from inception and design through deployment, operation, and refinement.
  • Support services before they go live through systemdesign consulting, capacity planning, and launch reviews.
  • Maintain live services by measuring and monitoringavailability, latency, and overall system health.
  • Support code deployments into multiple lowerenvironments, supporting current processes while automating wherever possible.
  • Support the application CI/CD pipeline for promotingsoftware into higher environments through validation and operational gating,and lead on DevOps automation and best practices.
  • Scale systems sustainably through automation, pushingfor changes that improve both reliability and velocity.
  • Collaborate with a global team spread across tech hubsin multiple geographies and time zones, sharing knowledge and explainingprocesses and procedures to others.
Requirements

Requirements

  • Strong hands-on experience with Linux.
  • Experience with monitoring tools such as Splunk,Dynatrace, or equivalent.
  • Working knowledge of ITIL/ITSM practices.
  • Strong troubleshooting skills across complex,multi-layered platforms.
  • Proficiency in SQL and PL/SQL.
  • Experience with Jenkins and CI/CD pipelines.
  • Scripting experience with Groovy, YAML, and Shell.
  • Experience with Git and Bitbucket.
  • Hands-on experience with Kubernetes and AWS.
  • Proven experience in production support leadership,including runbook and support model creation.
  • Experience defining monitoring and alerting strategies.
  • Experience with disaster recovery and resiliencyplanning.
  • Experience with deployment readiness validation andoperational process design.
  • Solid background in root cause analysis and problemmanagement.
  • A service ownership mindset with a focus on continuousimprovement and toil reduction.
  • Strong communication skills and the ability to shareknowledge across distributed teams.
  • Availability to participate in rotational on-callduties and occasional off-hours work.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Global SRE: Production Reliability & Automation
Global SRE: Production Reliability & Automation

Fulcrum Digital • Ciudad de México

Remote
MXN 600,000 - 1,000,000
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

On-site
MXN 1,049,685 - 1,399,580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

On-site
MXN 870,019 - 1,305,028
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer
Site Reliability Engineer

Infojini Inc • Mexico

Remote
MXN 1,200,000 - 1,800,000
Senior Software Engineer
Senior Software Engineer

Fulcrum Digital • Ciudad de México

Hybrid
MXN 600,000 - 900,000
Application support
Application support

Sequoia Connect LLC • Ciudad de México

Remote
MXN 446,000 - 781,000
Remote work
Senior Cloud/DevOps Engineer ID92208
Senior Cloud/DevOps Engineer ID92208

AgileEngine, LLC. • Región Centro

Hybrid
MXN 2,172,000 - 3,258,000
Professional growth
Competitive pay
Exciting projects
+1
Cloud Platform Technical Lead ID92209
Cloud Platform Technical Lead ID92209

AgileEngine, LLC. • Puebla de Zaragoza

On-site
MXN 2,172,000 - 3,258,000
Professional growth
Competitive USD-based pay
Exciting projects
+1
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Monterrey

Hybrid
MXN 1,262,000 - 1,983,000
Professional growth
Competitive USD pay
Exciting projects
+1
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Región Centro

Hybrid
MXN 420,000 - 700,000
Professional growth
Competitive USD-based compensation
Challenging projects with Fortune 500/