Site Reliability Engineer

Infojini Inc

México

Remote

MXN 1,200,000 - 1,800,000

Full time

10 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Infojini Inc. is seeking a Senior SRE to join our LATAM remote team, focusing on real-time incident response, automation, and reliability across multi-cloud environments. This role partners with US-based engineering teams to ensure availability and performance of critical systems.

Key responsibilities include incident command during outages, rapid RCA, and advancing automation to reduce toil, with a heavy emphasis on incident bridges and cross-functional collaboration.

Qualifications

  • Strong incident management and triage in production environments.
  • Solid Linux systems administration including performance and networking.
  • Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2).

Responsibilities

  • Act as first responder to alerts and production incidents, assessing severity and initiating mitigation actions.
  • Lead bridge calls as Incident Commander during major incidents.
  • Drive root cause isolation within 30 minutes for critical incidents when possible.
  • Communicate effectively across engineering, product, and leadership during high-pressure situations.

Skills

Incident management
Linux administration
AWS core services
GCP/Azure familiarity
CI/CD pipelines

Tools

GitHub
GitLab

Job description

Position: Sr SRE - Site Reliability Engineer - ( AWS Certified )

LATAM Remote

Below JD for SRE We have requirement for Senior SRE (x2) with 8 years of relevant experience.

Working Hours – 5 PM to 1 AM EST (Mon – Fri)

About the Role & Team

We are expanding our Site Reliability Engineering (SRE) organization. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems.

This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters—our SREs act decisively to prevent service degradation and protect the customer experience.

What You’ll Do
  • · Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions
  • · Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency
  • · Drive root cause isolation within 30 minutes for critical incidents whenever possible
  • · Communicate effectively across engineering, product, and leadership during high-pressure situations
  • · Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision‑making
Proactive Reliability Engineering

Identify patterns, trends, and signals to prevent incidents before they occur

  • · Continuously improve alert quality, reduce noise, and increase signal fidelity
  • · Partner with engineering teams to enhance system resilience and reliability
Automation & Toil Reduction

Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows

  • · Build and improve tooling across incident response, observability, and operations
  • · Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value
Platform & Systems Support
Troubleshoot across a hybrid ecosystem including:
  • · Cloud platforms (AWS, GCP, Azure)

Networking (connectivity, latency, DB access interruptions)

Diagnose and resolve issues across:
  • · CDN and traffic management layers (Akamai, waiting rooms – plus)
Required Technical Skills & Experience
Core Engineering & Operations

Strong experience in incident management and triage in production environments

  • · Proven ability to troubleshoot complex distributed systems under pressure
  • · Solid understanding of Linux systems administration (including performance, networking, NTP, etc.)
  • · Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)
  • · Familiarity with GCP and/or Azure environments
  • · Experience operating in multi-cloud and hybrid environments
Containers & Orchestration
  • · Understanding of containerized application architectures
DevOps & CI/CD

Strong knowledge of DevOps practices and CI/CD pipelines

Hands-on experience with:

  • · GitHub and/or GitLab
Working knowledge of:
  • · Java, Node.js, React-based applications
Understanding of database connectivity and dependencies across:
  • · Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required)
Networking

Strong foundational knowledge of:

  • · Load balancing and network troubleshooting
  • · Diagnosing connectivity issues between services and databases
Preferred Qualifications

Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications

Prior experience as an Incident Commander or similar leadership role during outages

Familiarity with Akamai CDN and traffic management tools

Experience in high-volume, high-availability production environments

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

On-site
MXN 1,049,685 - 1,399,580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

On-site
MXN 870,019 - 1,305,028
Professional growth
Competitive compensation
Exciting projects
+1
Lead SRE Engineer
Lead SRE Engineer

Cloudsufi • Región Centro

On-site
MXN 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

On-site
MXN 1,433,948 - 1,792,436
Junior SRE Engineer Guadalajara Mexico
Junior SRE Engineer Guadalajara Mexico

ReHire, LLC • Región Centro

On-site
MXN 1,080,000 - 1,621,000
Site Reliability Engineer
Site Reliability Engineer

Pyramid Consulting, Inc • Estado de México

On-site
MXN 334,800 - 558,000
SRE
SRE

Fulcrum Digital • Ciudad de México

Remote
MXN 600,000 - 1,000,000
Site Reliability Engineer
Site Reliability Engineer

Tata Consultancy Services • Ciudad de México

On-site
MXN 223,200 - 334,800
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Rosarito

On-site
MXN 1,531,000 - 2,212,000
Mentorship and TechTalks
Competitive USD-based compensation
Work on modern solutions
+1
Senior AWS Site Reliability Engineers - 2850
Senior AWS Site Reliability Engineers - 2850

Xideral • Región Centro

Hybrid
MXN 1,200,000 - 1,500,000
Premium Benefits
Performance bonuses
SGMM Medical insurance