AI Engineer (AWS, Splunk ITSI, Datadog, Python, New Relic, Azure, Dynatrace, ML, Service Now, Ansible, Rest API, ITIL)

EVYONIC SOLUTIONS PTE. LTD.

Singapore

On-site

SGD 200,000 - 260,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

EVYONIC SOLUTIONS PTE. LTD. is seeking a seasoned observability expert to design, implement and optimize enterprise-grade monitoring across cloud and on-prem environments.

You will lead integration of Splunk ITSI, Datadog, Dynatrace, New Relic, and AppDynamics, develop dashboards and SLIs/SLOs, and mentor teams in incident analysis and proactive monitoring to improve service reliability. Responsibilities span automation with Python and REST, collaboration with infra, app, cloud, and operations

Qualifications

  • 15+ years in AIOps, observability, infrastructure monitoring, or related IT ops.
  • Bachelor’s degree in Computer Science, Engineering, IT, or equivalent.
  • Hands-on with Splunk ITSI and enterprise observability platforms.
  • Experience with Datadog, Dynatrace, New Relic and/or AppDynamics.
  • Familiarity with ServiceNow, CMDB, CI relationships, and service mapping.
  • AIOps, ML, anomaly detection, event correlation, and predictive analytics.
  • Python, Linux, Ansible, and REST APIs for automation and integration.
  • Experience with AWS, Azure, and/or GCP cloud environments.
  • Strong ITIL, SRE, incident management knowledge and SLI/SLO concepts.
  • Analytical, troubleshooting skills using logs, metrics, events; strong mentoring.

Responsibilities

  • Design, implement, and optimize enterprise observability solutions using Splunk ITSI.
  • Architect and maintain observability solutions using Datadog for infrastructure, application, log, and service monitoring.
  • Implement and optimize Dynatrace for application performance monitoring, infrastructure visibility, and service health analysis.
  • Utilize New Relic for application monitoring, performance analysis, and proactive issue detection.
  • Implement AppDynamics for application performance monitoring, transaction analysis, and application health management.
  • Develop dashboards, KPIs, service views, alerts, and service-health monitoring for critical apps and infrastructure.
  • Analyze logs, events, metrics, and telemetry to investigate incidents, identify anomalies, and support root-cause analysis.
  • Implement intelligent event correlation, alert optimization, and noise-reduction techniques.
  • Apply AIOps, ML, and predictive analytics for anomaly detection and proactive incident management.
  • Integrate observability platforms with ServiceNow, CMDB, CI relationships, and service mapping.
  • Design observability solutions across AWS, Azure, and GCP cloud environments.
  • Establish and monitor SLI/SLO metrics to improve service reliability.
  • Automate monitoring and operations using Python, Linux, Ansible, and REST APIs.
  • Collaborate with infra, app, cloud, and operations teams to improve reliability and observability standards.

Skills

AIOps
Observability expertise
Stakeholder management
Mentoring
Analytical thinking
Incident management

Education

Bachelor’s degree in Computer Science or Engineering

Tools

Splunk ITSI
Datadog
Dynatrace
New Relic
AppDynamics
ServiceNow
CMDB
Python
Ansible
REST APIs
AWS
Azure
GCP
Linux

Job description

Responsibilities
  • Design, implement, and optimize enterprise observability solutions using Splunk ITSI.
  • Architect and maintain observability solutions using Datadog for infrastructure, application, log, and service monitoring.
  • Implement and optimize Dynatrace for application performance monitoring, infrastructure visibility, and service health analysis.
  • Utilize New Relic for application monitoring, performance analysis, and proactive issue detection.
  • Implement AppDynamics for application performance monitoring, transaction analysis, and application health management.
  • Develop and enhance dashboards, KPIs, service views, alerts, and service-health monitoring for critical applications and infrastructure.
  • Analyze logs, events, metrics, and telemetry to investigate incidents, identify anomalies, and support root-cause analysis.
  • Implement intelligent event correlation, alert optimization, and noise-reduction techniques to improve operational efficiency.
  • Apply AIOps, machine learning, and predictive analytics for anomaly detection, predictive monitoring, and proactive incident management.
  • Integrate observability platforms with ServiceNow, CMDB, CI relationships, and service mapping to improve service visibility and dependency mapping.
  • Design and implement observability solutions across AWS, Azure, and GCP cloud environments.
  • Establish and monitor SLI/SLO metrics to improve service reliability and operational performance.
  • Automate monitoring and operational processes using Python, Linux, Ansible, and REST APIs.
  • Collaborate with infrastructure, application, cloud, and operations teams to improve service reliability and observability standards.
Requirements
  • 15+ years of experience in AIOps, observability, infrastructure monitoring, systems engineering, or related IT operations roles.
  • Bachelor’s degree in computer science, Engineering, Information Technology, or equivalent.
  • Strong hands-on experience with Splunk ITSI and enterprise observability platforms.
  • Hands-on experience with Datadog, Dynatrace, New Relic, and/or AppDynamics.
  • Experience in ServiceNow, CMDB, CI relationships, and service mapping.
  • Experience with AIOps, machine learning, anomaly detection, event correlation, and predictive analytics.
  • Experience in Python, Linux, Ansible, and REST APIs for automation and integration.
  • Experience working with AWS, Azure, and/or GCP cloud environments.
  • Strong understanding of ITIL, SRE, incident management, and SLI/SLO concepts.
  • Strong analytical and troubleshooting skills using logs, metrics, and events.
  • Experience in designing and supporting enterprise-scale observability and monitoring platforms.
  • Strong stakeholder management, technical leadership, and mentoring skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer (AWS, Splunk ITSI, Datadog, Python, New Relic, Azure, Dynatrace, ML, Service Now, Ansible, Rest API, ITIL)
AI Engineer (AWS, Splunk ITSI, Datadog, Python, New Relic, Azure, Dynatrace, ML, Service Now, Ansible, Rest API, ITIL)

EXASOFT CONSULTING PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
AI Engineer (OPs Senior Level)
AI Engineer (OPs Senior Level)

antas pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
AI Ops Engineer
AI Ops Engineer

QUANTUM INFOTECH SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 170,000
AI Ops Engineer
AI Ops Engineer

TECHEMERGE SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
AI Ops Engineer (Splunk, ITSI, ServiceNow, CMBD, ML, Python, Linux, ITIL)
AI Ops Engineer (Splunk, ITSI, ServiceNow, CMBD, ML, Python, Linux, ITIL)

EVYONIC SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Ops Engineer (Splunk, ITSI, ServiceNow, CMBD, ML, Python, Linux, ITIL)
AI Ops Engineer (Splunk, ITSI, ServiceNow, CMBD, ML, Python, Linux, ITIL)

EXASOFT CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
AIOps Engineer
AIOps Engineer

Alphaeus Pte Ltd • Singapore

On-site
SGD 90,000 - 120,000
AI Ops Engineer (Splunk ITSI, ServiceNow CMDB, Machine Learning, Anomaly Detection, Event Correlation)
AI Ops Engineer (Splunk ITSI, ServiceNow CMDB, Machine Learning, Anomaly Detection, Event Correlation)

EVYONIC SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
AI Ops Engineer (Splunk ITSI, Glass Tables, CMDB, CI Synchronization, Event Correlation, ML-Based Predictive Analytics)
AI Ops Engineer (Splunk ITSI, Glass Tables, CMDB, CI Synchronization, Event Correlation, ML-Based Predictive Analytics)

EVYONIC SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 140,000 - 210,000
AI Ops engineer
AI Ops engineer

ANTAS PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000