Principal Service Reliability Engineer

Amadeus Hospitality

Manila

Hybrid

PHP 1,000,000 - 1,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amadeus Hospitality in Taguig, Metro Manila, is seeking a Principal Site Reliability Engineer to ensure the reliability, performance, and scalability of our mission-critical platforms. You will influence reliability strategies, participate in production incident responses, and guide capacity planning while collaborating with Development and Production Support teams to meet target SLOs.

The role is hybrid, requiring presence in the local office 2-3 days a week, and involves toil reduction,

Qualifications

  • Experience with cloud platforms, especially Azure.
  • Strong incident management and post-incident reviews.
  • Ability to drive reliability improvements across teams.

Responsibilities

  • Define SLI/SLOs and error budgets with product leads.
  • Lead toil reduction and automation projects.
  • Coordinate incident response and root cause analysis.
  • Develop playbooks for production issues.
  • Collaborate with Development and Operations to improve reliability.

Skills

Azure DevOps
Git/GitOps
Terraform
Docker
Kubernetes
PowerShell
Monitoring & Observability
Incident Response

Tools

Azure Runbooks
Grafana
Dynatrace
Splunk
Terraform

Job description

## Principal Service Reliability EngineerApplylocations: Taguig, Metro Manilatime type: Full timeposted on: Posted Todayjob requisition id: R35365**Job Title**Principal Service Reliability Engineer**Purpose of the role**The Principal Site Reliability Engineer will be responsible for ensuring the reliability, performance and scalability of our mission-critical platforms. In this role, you will be safeguarding operational excellence in the assigned product, influence reliability strategies, integral in production incident response, and helping to improve operational metrics. The role will be collaborating closely with different teams, such as Development and Production Support teams, to make sure target SLOs are met, making adjustments where needed and designing/developing code to facilitate meeting such targets. The role is also expected to work on toil reduction projects, handle capacity planning/tuning activities, revisiting existing SOPs and designing/developing code for performance improvements. This is a hybrid position and would require you to be in the local office 2-3 days a week.**In this role you'll:**- Define and track Service Level Indicators (SLIs), Objectives (SLOs), and Error Budgets in partnership with engineering and product leads- Collaborate with Operations and Development teams to drive service reliability, availability, and scalability - Drive and participate in toil reduction projects to minimize if not eliminate recurring manual activities performed by the team - Establish feedback loop with development teams for them to have visibility on the how stable and reliable their services are in client environments- Drive production incident response and lead root cause analysis and continuous improvement- Design/Develop operational improvement items with development teams working with them closely in prioritizing these improvements- Provide input on process improvements to Change, Release, and Incident Management - Create and implement support playbooks that resources can use as part of emergency response to production issues**About the ideal candidate**- Knowledgeable and experienced in utilizing different Azure resources such as VMs, Storage, Network, Functions, Logic Apps. App Services, AKS - Strong technical expertise on Azure DevOps, developing in git and working on gitops repo and build/release pipelines - Have hands-on experience in developing Azure Powershell scripts, Azure Runbooks, or any other infrastructure automation tools - Experienced with monitoring and logging tools (Grafana, Dynatrace, Splunk) - Proven ability to adapt to emerging cloud technologies and industry leading DevOps applications such as Terraform, Docker Containers, and Kubernetes- Knowledgeable in cloud implementation of Navitaire products across different cloud infrastructure models - Understands production environments and processes and ways on how they can be further optimized through various Azure features and other cloud technologies/services - Proven ability to drive problem solving efforts through effective issue analysis - Has the ability to lead efforts to implement infrastructure changes to increase environment stability and support scalability - Has the ability to drive collaborations with different Navitaire teams in enforcing environment standards and policies - Effectively works in a team environment and contributes in building capabilities of team members - Proficient in C# - Proven ability to work in a dynamic, fast-paced and multi-cultural environment - Willing to work on shifting schedules
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Service Reliability Engineer
Principal Service Reliability Engineer

Amadeus IT Group S.A. • Manila

Hybrid
PHP 1,200,000 - 1,800,000
Principal Service Reliability Engineer
Principal Service Reliability Engineer

Amadeus • Taguig

Hybrid
PHP 1,200,000 - 1,800,000
Senior Site Reliability Engineer – Azure & DevOps Lead
Senior Site Reliability Engineer – Azure & DevOps Lead

Amadeus • Taguig

Hybrid
PHP 1,200,000 - 1,800,000
Senior Site Reliability Engineer (Hybrid) - Azure & Cloud
Senior Site Reliability Engineer (Hybrid) - Azure & Cloud

Amadeus IT Group S.A. • Manila

Hybrid
PHP 1,200,000 - 1,800,000
Azure Site Reliability Engineer
Azure Site Reliability Engineer

GSS HR Solutions Private Limted • Quezon City

Hybrid
PHP 600,000 - 900,000
Site Reliability Engineer - BGC - Hybrid - Up to 180K
Site Reliability Engineer - BGC - Hybrid - Up to 180K

weSource Management Consultancy Firm • Taguig

Hybrid
PHP 150,000 - 180,000
Site Reliability Engineer | Hybrid - Centris/Makati
Site Reliability Engineer | Hybrid - Centris/Makati

TASQ • Makati

Hybrid
PHP 1,000,000 - 1,800,000
Principal Site Reliability Engineer (SRE)
Principal Site Reliability Engineer (SRE)

Lewis Personnel Management • Makati

Hybrid
PHP 1,800,000 - 3,200,000
Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month
Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month

AVENSYS CONSULTING INC. • Pasay

Hybrid
PHP 1,000,000 - 1,600,000
Azure Cloud Engineer
Azure Cloud Engineer

Accenture in the Philippines • Mandaluyong

On-site
PHP 600,000 - 1,200,000