Senior Site Reliability Engineer

AgileEngine

Mexico

Hybrid

MXN 1,525,000 - 2,034,000

Full time

12 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Mentorship
TechTalks
Growth roadmaps
Competitive USD-based compensation
Education budget
Fitness budget
Team activities
Exciting projects
Flextime
Hybrid work arrangement

Job summary

AgileEngine is seeking a Senior Site Reliability Engineer to ensure core system administration and operational stability across on-premise and SaaS-hosted environments. The role emphasizes Kubernetes cluster management, monitoring, and observability with Snowflake and OpenTelemetry.

You will participate in on-call rotations, incident response, and root-cause analysis, and automate infrastructure tasks using Python, Bash, or Go to keep systems healthy across multiple platforms.

Qualifications

  • 4+ years of infrastructure management experience.
  • Experience with Kubernetes.
  • Experience with monitoring, observability, and related SRE tasks.
  • Familiarity with Snowflake and OpenTelemetry ecosystems.
  • Strong background in Linux/Unix administration.
  • Proficiency in scripting languages (Bash, Python, or Go).

Responsibilities

  • Provide core system administration for enterprise infrastructure.
  • Maintain, scale, and ensure operational stability of on-premise and SaaS-hosted systems.
  • Manage and scale containerized environments using Kubernetes.
  • Participate in on-call rotations, incident response, and RCA.
  • Automate repetitive infrastructure tasks using scripting and IaC.

Skills

Kubernetes
Linux/Unix administration
Monitoring/Observability
OpenTelemetry
Snowflake
Bash
Python
Go
On-callIncidentManagement

Tools

Terraform
Ansible

Job description

We are looking for a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems, with a strong focus on Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry. You will participate in on-call rotations, incident response, and root-cause analysis, automate infrastructure tasks using Python, Bash, or Go, and ensure system health across ESM and ECP platform environments. Experience with service mesh architectures is highly valued for ECP-focused roles.

What you will do

  • Provide core system administration for the enterprise infrastructure.
  • Focus on the maintenance, scaling, and operational stability of various on-premise and SaaS-hosted systems.
  • Manage and scale containerized environments using Kubernetes.
  • For ECP-focused roles: Lean heavily into executing monitoring and observability tasks to ensure system health.
  • Participate in on-call rotations, incident response, and root-cause analysis (RCA).
  • Automate repetitive infrastructure tasks using scripting and infrastructure-as-code (IaC).

Must haves

  • 4+ years of infrastructure management experience.
  • Experience working with Kubernetes.
  • Experience with monitoring, observability, and related SRE tasks.
  • Familiarity with Snowflake and OpenTelemetry ecosystems.
  • Strong background in Linux/Unix administration.
  • Proficiency in scripting languages (e.g., Bash, Python, or Go).

Nice to haves

  • For ECP SREs: Experience with Service Mesh architectures is highly ideal.
  • Experience with Infrastructure as Code tools such as Terraform or Ansible.
  • Familiarity with cloud platforms (AWS, GCP, or Azure).

Perks and Benefits

  • Professional growth

Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps

  • Competitive compensation

We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities

  • A selection of exciting projects

Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands

  • Flextime

Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineer
Site Reliability Engineer

Infojini Inc • Mexico

Hybrid
MXN 900,000 - 1,300,000
Senior SRE: Kubernetes, Observability & IaC
Senior SRE: Kubernetes, Observability & IaC

AgileEngine • Mexico

Hybrid
MXN 1,525,000 - 2,034,000
Mentorship
TechTalks
Growth roadmaps
+7
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

N-iX • Mexico

Hybrid
PHP 5,552,000 - 8,020,000
Flexible working format
Competitive salary
Career growth
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

On-site
PHP 1,200,000 - 1,600,000
Medical / Health Insurance
Employee Assistance Programme
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Philtech Inc. • Taguig

On-site
PHP 670,000 - 1,339,000
Health insurance
Retirement plans
Career growth opportunities
Senior Kubernetes Engineer
Senior Kubernetes Engineer

OpsWerks • Mandaluyong

On-site
Senior Engineer - Site Reliability
Senior Engineer - Site Reliability

Dencom Consultancy and Manpower Services • Parañaque

On-site