Global SRE: Production Reliability & Automation

Fulcrum Digital

Ciudad de México

Remote

MXN 600,000 - 1,000,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Fulcrum Digital is seeking a Site Reliability Engineer to own the health, stability, and performance of production environments for a client engagement. You will define monitoring strategies, drive automation across deployments and operations, and collaborate with development teams to reduce incidents and improve resiliency.

Join a global team across time zones, implement scalable observability solutions, and ensure smooth code deployments with a strong service ownership mindset and proactive

Qualifications

  • Experience with monitoring tools such as Splunk or Dynatrace.
  • Working knowledge of ITIL/ITSM practices.
  • Strong troubleshooting across complex, multi-layered platforms.
  • Proficiency in SQL and PL/SQL.
  • Experience with Jenkins and CI/CD pipelines.
  • Hands-on scripting with Groovy, YAML, and Shell.
  • Experience with Git and Bitbucket.
  • Hands-on experience with Kubernetes and AWS.
  • Proven experience in production support leadership, including runbook and support model creation.
  • Experience defining monitoring and alerting strategies.
  • Experience with disaster recovery and resiliency planning.
  • Experience with deployment readiness validation and operational process design.
  • Solid background in root cause analysis and problem management.

Responsibilities

  • Plan, manage, and oversee all aspects of the production environment.
  • Define strategies for application performance monitoring and optimization in production.
  • Design, develop, and standardize monitoring and alerting mechanisms for supported applications.
  • Respond to incidents, improve the platform based on feedback, and measure the reduction of incidents over time.
  • Take a holistic approach to problem solving during production events, connecting the dots across the full technology stack to optimize mean time to recover (MTTR).
  • Analyze ITSM activities for the platform and provide a feedback loop to development teams on operational gaps or resiliency concerns.
  • Engage in and improve the whole lifecycle of services, from inception and design through deployment, operation, and refinement.
  • Support services before they go live through system design consulting, capacity planning, and launch reviews.
  • Maintain live services by measuring and monitoring availability, latency, and overall system health.
  • Support code deployments into multiple lower environments, supporting current processes while automating wherever possible.
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead on DevOps automation and best practices.
  • Scale systems sustainably through automation, pushing for changes that improve both reliability and velocity.
  • Collaborate with a global team spread across tech hubs in multiple geographies and time zones, sharing knowledge and explaining processes and procedures to others.

Skills

Linux
ITIL/ITSM
Troubleshooting
SQL/PLSQL
CI/CD (Jenkins)
Scripting (Groovy, YAML, Shell)
Git/Bitbucket
Kubernetes
AWS
Production support leadership
Monitoring strategy
Disaster recovery
DevOps automation
Communication

Tools

Splunk
Dynatrace
Jenkins
Kubernetes
AWS
Git
Bitbucket

Job description

Fulcrum Digital is seeking a Site Reliability Engineer to own the health, stability, and performance of production environments for a client engagement. You will define monitoring strategies, drive automation across deployments and operations, and collaborate with development teams to reduce incidents and improve resiliency.

Join a global team across time zones, implement scalable observability solutions, and ensure smooth code deployments with a strong service ownership mindset and proactive

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE
SRE

Fulcrum Digital • Ciudad de México

Remote
MXN 600,000 - 1,000,000
Site Reliability Engineer - Cloud, Kubernetes & Automation
Site Reliability Engineer - Cloud, Kubernetes & Automation

rctsglobal-com • Región Centro

On-site
MXN 885,000 - 1,062,000
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

On-site
MXN 1,049,685 - 1,399,580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

On-site
MXN 870,019 - 1,305,028
Professional growth
Competitive compensation
Exciting projects
+1
Senior SRE - Build Secure, Automated Cloud Reliability
Senior SRE - Build Secure, Automated Cloud Reliability

Precision Labs • Ciudad de México

Hybrid
MXN 1,200,000 - 2,000,000
Stock options
Annual bonus
Remote work
+1
Site Reliability Engineer
Site Reliability Engineer

Infojini Inc • Mexico

Remote
MXN 1,200,000 - 1,800,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

Cognizant • Mexico

On-site
MXN 900,000 - 1,500,000
Career growth opportunities
Competitive benefits
Inclusive culture
Lead Systems Reliability Engineer (Enterprise SRE)
Lead Systems Reliability Engineer (Enterprise SRE)

Talent Solutions ManpowerGroup • Ciudad de México

On-site
MXN 900,000 - 1,300,000
Senior SRE Lead: Reliability, Observability & AI Ops
Senior SRE Lead: Reliability, Observability & AI Ops

Cloudsufi • Región Centro

On-site
MXN 900,000 - 1,400,000
Lead SRE Engineer
Lead SRE Engineer

Cloudsufi • Región Centro

On-site
MXN 900,000 - 1,400,000