ART 1486 - Application Engineer (Site Reliability)

FPT Asia Pacific Pte Ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

FPT Asia Pacific Pte Ltd seeks a platform engineer to build and operate the AI platform that powers the firm’s modeling, workflow runtime, and experimentation environment.

You will own release scope, notes and rollback plans, lead automation, observability, and the pipeline that carries more services with less manual toil.

Qualifications

  • Degree in Computer Science/Information Technology or related fields.
  • Hands-on production Kubernetes experience, ideally AWS EKS, debugging pod failures and rollouts without a runbook.
  • Strong IaC skills with Terraform, including shared modules, state management and plan understanding.
  • Experience with cloud hyperscalers (AWS/Azure/GCP) in production, especially networking/VPC/IAM/Kubernetes/databases/object storage.
  • Release and change ownership: guiding scope, notes, coordination, change records and rollback plans; building pipelines.
  • CI/CD pipeline ownership; maintain and build pipelines, not just consume them.
  • Observability design judgement: know when to alert and when to avoid false positives.
  • Familiarity with SLOs and error budgets; ability to explain their impact on team decisions.
  • Scripting and automation for tooling beyond shell one-liners.
  • Incident handling experience: detect, diagnose, recover and post-incident changes.
  • Linux/container debugging: processes, networking, DNS, TLS, resource limits.
  • Positive learning mindset and strong communication; agile and adaptable.

Responsibilities

  • Build and operate the AI platform gateway to model providers, workflow runtime, portal and experimentation environment.
  • Own release management: scope, notes, coordination, change records and rollbacks; reduce manual toil as services grow.
  • Lead infrastructure automation, observability, and release discipline to prevent recurring problems.

Skills

Kubernetes (AWS EKS)
Terraform (IaC)
Cloud platforms (AWS/Azure/GCP)
CI/CD pipeline ownership
Observability design
SLOs and error budgets
Scripting and automation
Incident response
Linux/container debugging
Agile mindset
Communication skills

Education

Degree in Computer Science/Information Technology or related fields

Tools

Terraform
Kubernetes (AWS EKS)
Datadog
OpenSearch
Temporal
Kong/Envoy/API Gateway

Job description

Summary:

The successful candidate will build and operate the AI platform that the rest of the firm consumes, i.e. the gateway to model providers, the workflow and agent runtime, the enterprise AI portal, and the experimentation environment used by the organization's investment research.

The candidate is responsible to build the automation, infrastructure, observability and release discipline that stops the same class of problem recurring.

He or she will be the release manager, i.e. owning scope, coordination and change records for what goes live. The role will lead the platform carrying more services than it started with, and less manual effort holding it up.

Skillset (Must have)
  • Possess a degree in Computer Science/Information Technology or related fields.
  • Hands-on experience in Production Kubernetes, ideally Amazon EKS. Able to diagnose why a pod is failing, why a rollout is stuck, or why a node is under pressure, without a runbook.
  • Strong practical experience in Infrastructure as Code (IaC), ideally using Terraform. Experience in composing and applying shared modules, managing state across environments, reading a plan and knowing what it will actually do before applying it, and recovering when an apply fails part-way. The candidate does not need to have authored a reusable module library (that is owned centrally here), but need to be genuinely comfortable extending existing IaC.
  • Experience in Core Hyperscaler (AWS, Azure, or GCP) in production, ideally AWS - networking/VPC, IAM, managed Kubernetes, managed relational databases, object storage, and key management with a real understanding of least-privilege access.
  • Experience in release and change ownership. The candidate has personally owned releases into a controlled production environment: scope, notes, coordination with dependent teams, a change record and a rollback plan. This is a must-have requirement, and be able to build a pipeline and owns what goes through it.
  • Experience in CI/CD pipeline ownership, building and maintaining pipelines, and not only consuming them.
  • Experience in observability design judgement. Able to arg why one metric deserves an alert and another does not, and to describe both an alert that is fought to add and deliberately deleted.
  • Working knowledge of SLOs and error budgets. The candidate does not need to have owned an error-budget policy, but should have worked somewhere that ran on one and be able to explain what it changed about how the team made decisions.
  • Experience in scripting and automation. Able to write and maintain tooling others rely on, beyond shell one-liners.
  • Possess genuine incident experience. Able to walk through a production incident that personally responded to: what being seen first, how to narrow it down, what get wrong on the way, and what changes afterwards.
  • Experience in Linux and container debugging fundamentals, processes, networking, DNS, TLS, and resource limits.
  • Possess positive learning and collaborative mindset.
  • Strong analytical, problem-solving and troubleshooting skills.
  • Good written and verbal communication skills.
  • Agile, fast learner and able to adapt to changes.
Skillset (Good to have)
  • Datadog specifically: monitor design, SLOs, APM, and controlling alert noise.
  • Authoring Helm charts rather than only editing values files.
  • Data migration as part of a release - planning, reversibility, and verification.
  • Chaos or fault-injection experience - AWS Fault Injection Service, or comparable practice.
  • Running OpenSearch, or Temporal, as operational services.
  • Access and identity provisioning at enterprise scale (AD groups, entitlement workflows).
  • AWS cost visibility and optimization, including token-level cost attribution.
  • Prior experience in operating AI or LLM workloads, such as inference capacity, provider rate limits, unusually long request lifetimes.
  • Operating an API gateway at high throughput - Kong, Envoy, AWS API Gateway or similar.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ART 1486 - Application Engineer (Site Reliability)
ART 1486 - Application Engineer (Site Reliability)

FPT Asia Pacific • Singapore

On-site
SGD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Application Engineers (AI Platform)
Site Reliability Application Engineers (AI Platform)

ASTEK SINGAPORE INNOVATION TECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 90,000 - 120,000
AI Engineer - Agentic & GenAI Systems
AI Engineer - Agentic & GenAI Systems

JOY CONSULTING PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
DevSecOps Engineer - #1634
DevSecOps Engineer - #1634

JOBSTER PRIVATE LTD. • Singapore

On-site
SGD 140,000 - 210,000
AI Engineer
AI Engineer

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
#EG AI Engineer
#EG AI Engineer

NCS Group • Singapore

On-site
SGD 80,000 - 120,000
Forward Deployed Engineer
Forward Deployed Engineer

TECHKNOWLEDGEY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Forward Deployed Engineer
Forward Deployed Engineer

TechKnowledgey Pte Ltd • Singapore

On-site
SGD 150,000 - 190,000