Graphite - Site Reliability Engineer (SRE)

rctsglobal-com

Región Centro

On-site

MXN 885,000 - 1,062,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

rctsglobal-com is seeking a hands-on Site Reliability Engineer (SRE) to ensure the stability, performance, and scalability of production services. You will bridge development and operations, using code to automate infrastructure and reduce toil across cloud environments.

Responsibilities include monitoring design, incident response, production debugging, and continuous improvement through blameless post-mortems. Proficiency in Python/Go, Terraform or Pulumi, and cloud platforms is required.

Qualifications

  • Proven experience as an SRE, DevOps Engineer, or similar role.
  • Expertise with Kubernetes in a production environment.
  • Strong scripting/programming skills (Python, Go, Bash).
  • Deep understanding of monitoring, logging, and alerting best practices.
  • Solid experience with at least one major Cloud provider (AWS, GCP, or Azure).
  • Experience with Infrastructure as Code tools like Terraform or Pulumi is a plus.

Responsibilities

  • Design, implement, and maintain robust monitoring and alerting systems to provide visibility into application performance and infrastructure health.
  • Build, provision, and maintain core infrastructure with emphasis on cloud environments and Kubernetes clusters.
  • Write and maintain scripts and automation workflows to streamline deployment, scaling, and operational tasks.
  • Provide hands-on, real-time incident response and participate in on-call rotation.
  • Debug and troubleshoot complex production problems across the stack.
  • Conduct blameless post-mortems and implement long-term improvements to prevent recurrence.

Skills

Python
Go
Bash
SRE/DevOps

Tools

Kubernetes
Terraform
Pulumi
AWS
GCP
Azure

Job description

Site Reliability Engineer (SRE)
Overview

We're looking for a passionate and hands‑on Site Reliability Engineer (SRE) to join our team. This role is critical for ensuring the stability, performance, and scalability of our production services. You'll be the bridge between development and operations, with a strong focus on using code to manage infrastructure and eliminate toil.

Key Responsibilities
  • Monitoring and Alerting: Design, implement, and maintain robust monitoring and alerting systems (e.g., GCP Monitoring, Prometheus, Grafana, Traces, Logs) to provide visibility into application performance and infrastructure health.
  • Infrastructure Management: Build, provision, and maintain our core infrastructure, with a strong emphasis on Cloud environments and Kubernetes clusters.
  • Automation and Tooling: Write and maintain scripts and automation workflows (e.g., Python, Bash, TypeScript (Pulumi)) to streamline deployment, scaling, and operational tasks, embracing the philosophy of "automating everything."
  • Incident Response: Provide hands‑on, real‑time incident response and participate in an on‑call rotation to quickly mitigate service disruptions and restore functionality.
  • Production Debugging: Deeply debug and troubleshoot complex production problems across the entire stack, from network issues to application code defects.
  • Process Improvement: Conduct blameless post‑mortems for major incidents, implementing long‑term solutions to prevent recurrence and continuously improve service reliability.
Qualifications
  • Proven experience as an SRE, DevOps Engineer, or similar role.
  • Expertise in managing and scaling Kubernetes in a production environment.
  • Strong proficiency in a scripting or programming language (e.g., Python, Go, Bash).
  • Deep understanding of monitoring, logging, and alerting best practices.
  • Solid experience with at least one major Cloud provider (AWS, GCP, or Azure).
  • Experience with Infrastructure as Code (IaC) tools like Terraform or Pulumi is a plus.
What You'll Bring

A proactive, data-driven approach to reliability and a passion for managing complex systems at scale.

Compensation

The base pay range for this role is $50,000 – $60,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

On-site
MXN 1,049,685 - 1,399,580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

On-site
MXN 870,019 - 1,305,028
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer - Cloud, Kubernetes & Automation
Site Reliability Engineer - Cloud, Kubernetes & Automation

rctsglobal-com • Región Centro

On-site
MXN 885,000 - 1,062,000
Site Reliability Engineer
Site Reliability Engineer

Infojini Inc • Mexico

Remote
MXN 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

On-site
MXN 1,433,948 - 1,792,436
Site Reliability Engineer
Site Reliability Engineer

Pyramid Consulting, Inc • Estado de México

On-site
MXN 334,800 - 558,000
Remote Site Reliability Engineer: Cloud & Kubernetes
Remote Site Reliability Engineer: Cloud & Kubernetes

itD • Estado de México

Remote
MXN 1,438,072 - 1,977,350
Comprehensive medical benefits
401K with matching
Paid holidays
+1
Site Reliability Engineer
Site Reliability Engineer

Tata Consultancy Services • Ciudad de México

On-site
MXN 223,200 - 334,800
Lead SRE Engineer
Lead SRE Engineer

Cloudsufi • Región Centro

On-site
MXN 900,000 - 1,400,000
Junior SRE Engineer Guadalajara Mexico
Junior SRE Engineer Guadalajara Mexico

ReHire, LLC • Región Centro

On-site
MXN 1,080,000 - 1,621,000