Site Reliability Engineer

Infojini Inc

Mexico

Hybrid

MXN 900,000 - 1,300,000

Full time

21 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infojini Inc is expanding its Site Reliability Engineering team to ensure reliability, availability, and performance of production systems. You will be part of offshore SRE squads collaborating with US-based engineers to respond to incidents in real time and drive proactive reliability improvements across cloud platforms.

The role emphasizes real-time incident response, automation, and strong collaboration across engineering, product, and leadership to protect customer experience in a

Qualifications

  • Strong incident management in production environments.
  • Troubleshoot complex distributed systems under pressure.
  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Hands-on with Linux systems administration and networking fundamentals.
  • Familiarity with CI/CD pipelines and DevOps practices.

Responsibilities

  • Act as first responder to alerts and production incidents, assessing severity and initiating mitigation actions.
  • Lead major incidents as Incident Commander, guiding bridge calls with clarity and urgency.
  • Drive root cause analysis within 30 minutes for critical incidents when possible.
  • Communicate across engineering, product, and leadership during high-pressure situations.
  • Reduce toil by automating operational tasks and building tooling for incident response.

Skills

Incident management
Linux administration
AWS
GCP
Azure
Multi-cloud
CI/CD
GitHub/GitLab
Containerization
Networking
Databases
Java
Node.js
React
Load balancing
AI-assisted tooling

Job description

Working Hours: 5 PM to 1 AM EST (Mon – Fri)

About the Role & Team

We are expanding our Site Reliability Engineering (SRE) organization. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems.

This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters—our SREs act decisively to prevent service degradation and protect the customer experience.

What You’ll Do
  • Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions
  • Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency
  • Drive root cause isolation within 30 minutes for critical incidents whenever possible
  • Communicate effectively across engineering, product, and leadership during high-pressure situations
  • Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision-making
Proactive Reliability Engineering
  • Identify patterns, trends, and signals to prevent incidents before they occur
  • Continuously improve alert quality, reduce noise, and increase signal fidelity
  • Partner with engineering teams to enhance system resilience and reliability
Automation & Toil Reduction
  • Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows
  • Build and improve tooling across incident response, observability, and operations
  • Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value
Platform & Systems Support

Troubleshoot across a hybrid ecosystem including:

  • Cloud platforms (AWS, GCP, Azure)
Diagnose and resolve issues across:
  • Networking (connectivity, latency, DB access interruptions)
  • CDN and traffic management layers (Akamai, waiting rooms – plus)
Required Technical Skills & Experience
Core Engineering & Operations
  • Strong experience in incident management and triage in production environments
  • Proven ability to troubleshoot complex distributed systems under pressure
  • Solid understanding of Linux systems administration (including performance, networking, NTP, etc.)
  • Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)
  • Familiarity with GCP and/or Azure environments
  • Experience operating in multi-cloud and hybrid environments
Containers & Orchestration
  • Understanding of containerized application architectures
DevOps & CI/CD

Strong knowledge of DevOps practices and CI/CD pipelines

Hands-on experience with:

  • GitHub and/or GitLab
Working knowledge of:
  • Java, Node.js, React-based applications
Understanding of database connectivity and dependencies across:
  • Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required)
Networking

Strong foundational knowledge of:

  • Load balancing and network troubleshooting
  • Diagnosing connectivity issues between services and databases
Preferred Qualifications
  • Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications
  • Prior experience as an Incident Commander or similar leadership role during outages
  • Familiarity with Akamai CDN and traffic management tools
  • Experience in high-volume, high-availability production environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineer
Site Reliability Engineer

Procter & Gamble • Manila

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Alsons/AWS Information Systems Inc. • Cebu City

Hybrid
PHP 600,000 - 1,000,000
JD IC4 - Sr Infra Engineer - SRE
JD IC4 - Sr Infra Engineer - SRE

Spin Careers • Philippines

Remote
PHP 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

Philtech Inc. • Taguig

On-site
PHP 670,000 - 1,339,000
Health insurance
Retirement plans
Career growth opportunities
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability / Cloud Platform Engineer
Site Reliability / Cloud Platform Engineer

Global Recruitment and Consultancy OPC • Cebu City

On-site
PHP 1,200,000 - 2,400,000
Senior Engineer - Site Reliability
Senior Engineer - Site Reliability

Dencom Consultancy and Manpower Services • Parañaque

On-site
Production Engineer
Production Engineer

ECLARO • Taguig

On-site
PHP 900,000 - 1,200,000