Site Reliability Engineer

AgileEngine

Mexico

Hybrid

PHP 8,734,000 - 13,100,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Professional growth
Competitive USD-based pay
Exciting projects
Flextime

Job summary

AgileEngine is seeking a DevOps / Site Reliability Engineer to sustain operational resilience across Azure, AWS, and GCP in a 24x7 setting. You will bridge platform engineering with incident command, own remediation, and shape playbooks to guide rapid response.

You will lead major-incident calls, design automated runbooks, and drive post-incident reviews, while advancing security and compliance through IaC, CI/CD, and CSPM tooling.

Qualifications

  • 5+ years of experience required.
  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
  • Strong hands-on experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
  • Senior-level incident-command experience in a 24x7 production environment.
  • Proven ability to drive remediation follow-up and coordinate across teams.
  • Experience drafting incident notifications for technical and executive audiences.
  • Experience authoring incident-management playbooks and escalation procedures.
  • Autonomy in managing complex operations.

Responsibilities

  • Scale and maintain stability across multi-cloud environments (Azure, AWS, GCP).
  • Engineer IaC-based security policies and baselines to prevent misconfigurations.
  • Design and maintain enterprise CI/CD pipelines for ingestion and deployment.
  • Respond to continuous monitoring alerts using CSPM tools like Wiz.
  • Act as Incident Commander for major incidents in a 24x7 setting.
  • Own post-incident remediation and drive systemic fixes with cross-team accountability.
  • Draft timely incident notifications for technical and executive audiences.
  • Develop and socialize incident-management playbooks and runbooks.

Skills

Kubernetes
Terraform
CI/CD
Python/Go
Incident Command
Multi-cloud
Zero-trust
Wiz
PCI-DSS
SOC2

Tools

Wiz

Job description

We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.

What you will do
  • Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
  • Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
  • Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
  • Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
  • Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
  • Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
  • Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
  • Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.
Must haves
  • 5+ years of experience.
  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
  • Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
  • Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
  • Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
  • Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
  • Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
  • Fully autonomous.
  • Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
  • Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
  • Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
Nice to haves
  • PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
  • ServiceNow — familiarity with incident, problem, and change management workflows and reporting.
Perks and Benefits
  • Professional growth

Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps

  • Competitive compensation

We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities

  • A selection of exciting projects

Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands

  • Flextime

Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Multi-Cloud Incident Command
Senior SRE: Multi-Cloud Incident Command

AgileEngine • Mexico

Hybrid
PHP 8,734,000 - 13,100,000
Professional growth
Competitive USD-based pay
Exciting projects
+1
Site Reliability Engineer
Site Reliability Engineer

IDEMIA PHILIPPINES INC. • Philippines

On-site
PHP 900,000 - 1,350,000
Site Reliability / Cloud Platform Engineer
Site Reliability / Cloud Platform Engineer

Global Recruitment and Consultancy OPC • Cebu City

On-site
PHP 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Philtech Inc. • Taguig

On-site
PHP 670,000 - 1,339,000
Health insurance
Retirement plans
Career growth opportunities
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

On-site
PHP 1,200,000 - 1,600,000
Medical / Health Insurance
Employee Assistance Programme
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineers
Site Reliability Engineers

Trinity Workforce Solutions, Inc. • Makati

On-site
Senior IT Engineer
Senior IT Engineer

OpsWerks • Mandaluyong

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila, Hinoba-an

On-site
PHP 1,800,000 - 3,000,000