G13 - Operations Support Engineer

FPT Asia Pacific

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

FPT Asia Pacific is seeking an Operations Support Engineer to design and own the service observability usage model, ensuring metrics, logs and traces flow into Elastic Cloud. You will build proactive alerting, incident response playbooks and RCA workflows to reduce MTTR and improve reliability.

The role requires strong SRE experience, hands-on AWS/Kubernetes skills, IaC/GitOps, and scripting for ops automation, with focus on secure supply chain controls and policy enforcement.

Qualifications

  • 4+ years in SRE/Production Ops for SaaS or high-throughput services.
  • Working knowledge of AWS and Kubernetes for collaboration with platform owners.
  • Familiarity with Infrastructure as Code and GitOps (Terraform, Argo).
  • Observability implementation with Elastic Cloud; Prometheus/OpenTelemetry concepts.
  • Proven on-call and incident management experience; MTTR reduction.
  • Scripting in Python, Bash, or Go for ops tooling.
  • Security/compliance awareness: vulnerability management and supply chain controls.
  • Clear, concise communication of operational risk to stakeholders.

Responsibilities

  • Design & own service observability usage model; ensure metrics/logs/traces flow into Elastic Cloud.
  • Build proactive alerting and incident response playbooks; drive RCA and remediation tracking.
  • Optimize service performance toward latency and throughput targets.
  • Implement secure supply chain and runtime controls (image scanning, SBOM, TLS/mTLS).
  • Curate runbooks, dashboards, and production readiness checklists.
  • Support compliance & audit evidence collection via automated evidence capture.
  • Introduce drift detection & policy-as-code guardrails (OPA/Kyverno).
  • Mentor engineers on production readiness and on-call playbooks.
  • Participate in equitable on-call rotation with sustainable alert volumes.

Skills

SRE
AWS
Kubernetes
Terraform
GitOps
Observability
Incident management
Scripting

Tools

Elastic Cloud
Prometheus
OpenTelemetry
Grafana
CloudWatch
Argo
OPA
Kyverno

Job description

About the job G13 - Operations Support Engineer

Responsibilities:

  • Design & own service observability usage model: ensure all service metrics, logs, traces flow into Elastic Cloud (authoritative); maintain dashboards & SLOs; evaluate pragmatic use of CloudWatch, AWS Managed Prometheus / Grafana for supplemental or fallback views.
  • Build proactive, noise-reduced alerting and incident response playbooks; drive post-incident RCA & remediation tracking (closure SLA).
  • Optimize service performance (profiling, caching layers, autoscaling heuristics, concurrency tuning) meeting latency & throughput targets.
  • Implement secure supply chain & runtime controls (image scanning, SBOM consumption, secrets management, TLS / mTLS) leveraging shared platform tooling.
  • Curate operational runbooks, golden dashboards, reliability readiness + production readiness checklists.
  • Support compliance & audit evidence collection (access logs, config lineage, change histories) via automated evidence capture fed into Elastic.
  • Introduce configuration drift detection & policy-as-code guardrails (OPA / Kyverno) at the workload / namespace layer to enforce baseline controls.
  • Mentor engineers on production readiness, observability patterns, and operational excellence; evolve on-call playbooks.
  • Participate in (and improve) an equitable on-call rotation focusing on sustainable alert volumes & burnout prevention.

Requirements

  • 4+ years (or equivalent impact) in SRE / Production Ops / Platform / Reliability for SaaS or high-throughput services.
  • Working knowledge of AWS & Kubernetes (deployment, troubleshooting, networking concepts) sufficient to collaborate effectively with platform owners (not necessarily owning cluster upgrade orchestration).
  • Familiarity with Infrastructure as Code & GitOps (Terraform, Argo, etc.) to consume modules, review changes, and enforce policy.
  • Observability implementation & usage (metrics, logs, traces, profiling) with Elastic Cloud; understanding of Prometheus / OpenTelemetry concepts.
  • Proven on-call & incident management experience (triage, MTTR reduction, RCA authorship).
  • Scripting / automation in Python, Bash, or Go for ops tooling.
  • Security & compliance aware: vulnerability management, image scanning, supply chain controls.
  • Clear, concise communication of operational risk & trade-offs to technical + non-technical stakeholders.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

U3 SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Cloud Operations Engineer – Infrastructure
Cloud Operations Engineer – Infrastructure

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 110,000 - 170,000
Cloud Operations Engineer
Cloud Operations Engineer

Hamilton Barnes ? • Singapore

On-site
SGD 70,000 - 110,000
Platform Operations Engineer (Cloud Infrastructure)
Platform Operations Engineer (Cloud Infrastructure)

TEKISHUB CONSULTING SERVICES PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Software & Applications Manager (Technical Lead/Supervisor)
Software & Applications Manager (Technical Lead/Supervisor)

optimum solutions (singapore) pte ltd • Singapore

On-site
SGD 120,000 - 180,000
Cloud Engineer (Monitoring / AWS)
Cloud Engineer (Monitoring / AWS)

SEARCH INDEX PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )
Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )

NEPTUNEZ SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Head of Site Reliability Engineering (SRE) & Information Security
Head of Site Reliability Engineering (SRE) & Information Security

Kristal Advisors (Sg) Pte. Ltd. • Singapore

On-site
SGD 180,000 - 260,000
IT Infrastructure Engineer
IT Infrastructure Engineer

GMP RECRUITMENT SERVICES (S) PTE LTD • Singapore

On-site
SGD 60,000 - 110,000
SRE & Observability Engineer — Reliability & Automation
SRE & Observability Engineer — Reliability & Automation

FPT Asia Pacific • Singapore

On-site
SGD 120,000 - 180,000