Data Engineer | Site Reliability & Data Platform

NEXBRIDGE RECRUITMENT PTE. LTD.

Singapore

On-site

SGD 90,000 - 130,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NEXBRIDGE RECRUITMENT PTE. LTD. is seeking a skilled Observability/DevOps professional to design, implement and maintain monitoring across applications, infrastructure, cloud and hybrid environments.

You will work with metrics, logs, traces and dashboards to provide visibility into system health and performance. You will collaborate with engineering, infrastructure and security teams to improve reliability, runbooks and incident response while applying best practices in OpenTelemetry, Grafana

Qualifications

  • Minimum 3 years of experience in Observability, SRE, Platform Engineering, Infrastructure Engineering, DevOps or related areas.
  • Hands-on experience with production monitoring and observability environments.
  • Good understanding of metrics, logging, tracing, dashboards and alerting.
  • Experience with AWS and/or cloud environments.
  • Familiarity with Docker, CI/CD, Terraform/OpenTofu, Ansible or similar technologies.
  • Experience working with hybrid or distributed environments.
  • Strong troubleshooting and problem-solving skills.

Responsibilities

  • Design and maintain observability solutions across applications, infrastructure, cloud and hybrid environments.
  • Work with metrics, logs, events and traces to provide visibility into system health and performance.
  • Develop and maintain dashboards, monitoring solutions, alerts and service health indicators.
  • Support application and infrastructure teams with instrumentation and telemetry integration.
  • Work with technologies such as OpenTelemetry, Grafana, Prometheus, Dynatrace, Elastic or equivalent platforms.
  • Support incident investigation, troubleshooting and root-cause analysis.
  • Identify monitoring gaps and recommend improvements to system reliability and performance.
  • Work with engineering, infrastructure, network, security and platform teams on operational improvements.
  • Maintain technical documentation, operational procedures and runbooks.

Skills

Observability
SRE
Platform Engineering
Infrastructure Engineering
DevOps
Troubleshooting
Cloud environments

Tools

OpenTelemetry
Grafana
Prometheus
Dynatrace
Elastic
Docker
CI/CD
Terraform/OpenTofu
Ansible

Job description

Job Responsibilities
  • Design and maintain observability solutions across applications, infrastructure, cloud and hybrid environments.
  • Work with metrics, logs, events and traces to provide visibility into system health and performance.
  • Develop and maintain dashboards, monitoring solutions, alerts and service health indicators.
  • Support application and infrastructure teams with instrumentation and telemetry integration.
  • Work with technologies such as OpenTelemetry, Grafana, Prometheus, Dynatrace, Elastic or equivalent platforms.
  • Support incident investigation, troubleshooting and root-cause analysis.
  • Identify monitoring gaps and recommend improvements to system reliability and performance.
  • Work with engineering, infrastructure, network, security and platform teams on operational improvements.
  • Maintain technical documentation, operational procedures and runbooks.
Job Requirement
  • Minimum 3 years of experience in Observability, SRE, Platform Engineering, Infrastructure Engineering, DevOps or related areas.
  • Hands-on experience with production monitoring and observability environments.
  • Good understanding of metrics, logging, tracing, dashboards and alerting.
  • Experience with AWS and/or cloud environments.
  • Familiarity with Docker, CI/CD, Terraform/OpenTofu, Ansible or similar technologies.
  • Experience working with hybrid or distributed environments.
  • Strong troubleshooting and problem-solving skills.
Good to Have
  • Experience with OpenTelemetry at scale.
  • AWS or Azure certification.
  • Experience with large-scale or distributed environments.
  • Familiarity with SRE practices, SLOs, incident response and reliability engineering.

We regret that only shortlisted candidates will be notified.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer – Observability / SRE
Data Engineer – Observability / SRE

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 120,000
Data Engineer - Observability / SRE
Data Engineer - Observability / SRE

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer, Observability ( Contract )
Site Reliability Engineer, Observability ( Contract )

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Observability Engineer – SRE / Cloud Data
Observability Engineer – SRE / Cloud Data

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 80,000 - 140,000
Observability Engineer - SRE / Cloud Data
Observability Engineer - SRE / Cloud Data

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Assistant Lead Engineer - Observability Dynatrace (Ops Response)
Assistant Lead Engineer - Observability Dynatrace (Ops Response)

Synapxe • Singapore

On-site
SGD 120,000 - 180,000
Data Platform Engineer
Data Platform Engineer

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 120,000
Data Platform Engineer - Cloud & Data Infrastructure - Contract
Data Platform Engineer - Cloud & Data Infrastructure - Contract

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Systems Engineer - Observability (Infra Deliivery)
Systems Engineer - Observability (Infra Deliivery)

Synapxe • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

NTT Data Singapore • Singapore

On-site
SGD 110,000 - 190,000