Senior ITSMA Observability Engineer

HedgeServ

Manila

On-site

PHP 1,200,000 - 1,800,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

The Senior ITSMA Observability Engineer at HedgeServ designs and develops the Elastic and Prometheus Stack, along with AWS Observability tools, to monitor and manage HedgeServ’s critical applications and infrastructure. You will operate and design the monitoring tools portfolio, including alerting, dashboards, and the end-to-end framework supporting HedgeServ’s systems.

You will collaborate with the ITSMA Monitoring and Analytics Team to build secure, scalable solutions using Elastic Cloud Stack

Qualifications

  • Experience with Elastic Cloud and AWS Managed Prometheus.
  • Knowledge of installation, system tasks, data collection, network troubleshooting, data pipelines, and cluster administration.
  • Proficient in Python, Bash, PowerShell, Painless, and other scripting languages.
  • Extensive ELK Stack expertise: Elasticsearch, Logstash, Kibana, Beats, APM, X-Pack, REST API integration.
  • Skilled in evaluating and tuning Elastic clusters, configurations, indexing, search performance, security, and administration.
  • Proficient with Prometheus, Grafana, AWS observability tools, and their performance, security, and management.
  • Experienced with security integrations (Windows SAML, LDAP, Kerberos) in Elasticsearch.
  • Adept with AWS services: CloudWatch, CloudTrail, Kubernetes, Docker, Lambda.
  • Integrated Elastic alerting with third-party ticketing tools.
  • Experience in implementing observability AI agents and frameworks for automated analysis, incident detection, and proactive resolution across complex systems.

Responsibilities

  • Design, build, secure, maintain, optimize, and document solutions using Elastic Cloud Stack and AWS-managed Prometheus.
  • Collaborate with the ITSMA Monitoring and Analytics Team to implement observability tooling.
  • Develop alerts via Elastic Watcher and Kibana Alerts integrated with ticketing tools and MS Teams.
  • Lead IT infrastructure monitoring projects and vendor management.
  • Participate in agile sprint meetings and share progress to align with organizational requirements.

Skills

Elastic Stack
Prometheus
Python scripting
Terraform
Ansible
Kibana
ETL pipelines

Education

Technical Degree in Information Technology

Tools

Elasticsearch
Logstash
Grafana
AWS CloudWatch
OTEL (OpenTelemetry)
Kubernetes
Terraform

Job description

The Senior ITSMA Observability Engineer is responsible for the design and development of the Elastic and Prometheus Stack, as well as, AWS Observability tools that monitor and manage critical applications and infrastructure at HedgeServ. As an important member of the ITSMA Monitoring and Analytics Team, the Senior Engineer will be responsible for the operation and design of the portfolio of tools, which include alerting mechanisms and escalation, dashboards, and the overall framework to support the management of HedgeServ’s infrastructure, systems, and applications. Additionally, this role entails leading IT infrastructure monitoring projects and vendor management and handling daily operations with SME (Subject Matter Expert) escalation support as needed. The successful applicant should possess the ability to collaborate with various IT teams to gather requirements and develop solutions by means of existing monitoring capabilities or customized monitors (scripts).

Role Responsibilities

The Senior ITSMA Observability Engineer will collaborate with the ITSMA Monitoring and Analytics Team to design, build, secure, maintain, optimize, and document solutions utilizing Elastic Cloud Stack and AWS-managed Prometheus.

  • Proficiency with Elasticsearch, Logstash, Kibana, Beats, APM with X-Pack, Prometheus, Grafana, AWS CloudWatch, and other observability tools.
  • Experience with OTEL Collectors.
  • Engage closely with application owners, engineers, and development teams to evaluate requirements, architect, and support an Elasticsearch Stack solution, as well as structure queries to enhance system performance and efficiency.
  • Design and configure ETL data pipelines using Elastic Common Schema for onboarding application logs and metrics.
  • Configure index templates and manage data lifecycle (ILM) for effective data retention.
  • Develop Ansible playbooks for automated deployment of Beat agents across on-premises and AWS systems; utilize Terraform for safe management of production infrastructure, employing methodologies such as Infrastructure as Code within AWS environments.
  • Create Elastic alerting solutions via Watcher and Kibana Alerts integrated with existing ticketing tools and MS Teams.
  • Develop Machine Learning jobs to dynamically monitor and provide alerts based on specific metrics and KPIs.
  • Build Elastic and AWS observability AI solutions that enable infrastructure engineering and operations teams to address production issues efficiently.
  • Adhere to lifecycle processes for transitioning solutions from Development to QA to Production.
  • Actively participate in collaborative group sessions, attend agile sprint daily meetings, and share progress to ensure solution development aligns with organizational requirements.
Pre-Requisite Knowledge, Skills and Experience
  • Technical Degree in Information Technology
  • Experience with Elastic Cloud and AWS Managed Prometheus
  • Knowledge of installation, system tasks, data collection, network troubleshooting, data pipelines, and cluster administration
  • Proficient in Python, Bash, PowerShell, Painless, and other scripting languages
  • Extensive ELK Stack expertise, including Elasticsearch, Logstash, Kibana, Beats, Machine Learning, APM, X-Pack, and REST API integration
  • Skilled in evaluating and tuning Elastic clusters, configurations, indexing, search performance, security, and administration
  • Proficient with Prometheus, Grafana, AWS observability tools, and their performance, security, and management
  • Experienced with security integrations (Windows SAML, LDAP, Kerberos) in Elasticsearch
  • Adept with AWS services: CloudWatch, CloudTrail, Kubernetes, Docker, Lambda
  • Integrated Elastic alerting with third-party ticketing tools
  • Experienced in implementing and integrating observability AI agents and frameworks for automated analysis, incident detection, and proactive resolution across complex systems
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Observability Engineer (Elastic & AWS)
Senior Observability Engineer (Elastic & AWS)

HedgeServ • Manila

On-site
PHP 1,200,000 - 1,800,000
Senior ITSMA Observability Engineer
Senior ITSMA Observability Engineer

HedgeServ Corporation • Manila

On-site
PHP 4,299,754 - 5,528,255
Fully paid comprehensive health benefits
Remote and hybrid working arrangements
Senior Observability Engineer - Elastic & AWS, Remote
Senior Observability Engineer - Elastic & AWS, Remote

HedgeServ Corporation, HedgeServ Limited • Manila

Hybrid
PHP 4,299,754 - 5,528,255
Fully paid comprehensive health benefits
Remote and hybrid working arrangements
Senior IT Operations Engineer
Senior IT Operations Engineer

ECLARO • Quezon City

On-site
PHP 450,000 - 900,000
Monitoring, Observability and Event
Monitoring, Observability and Event

Gratitude Philippines • Quezon City

Hybrid
PHP 1,116,000 - 1,205,000
Information Technology Operations Engineer
Information Technology Operations Engineer

ECLARO • Philippines

On-site
PHP 600,000 - 900,000
Lead Security Engineer
Lead Security Engineer

Hammerjack Pty Ltd • Philippines

On-site
PHP 700,000 - 1,100,000
Lead Security Engineer
Lead Security Engineer

Vaco by Highspring • Philippines

On-site
PHP 800,000 - 1,200,000
Site Reliability Engineer | Hybrid - Centris/Makati
Site Reliability Engineer | Hybrid - Centris/Makati

TASQ • Makati

Hybrid
PHP 900,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia Natural Language Solutions Ua Ltd • Philippines

On-site
PHP 1,000,000 - 1,800,000
Fixed compensation
Long-term vacation
Career development courses
+1