Site Reliability Engineer (AI Platform)

Allegis Group Singapore Pte Ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Allegis Group Singapore Pte Ltd is seeking a Site Reliability Engineer to join a high-performing AI Platform team responsible for building and operating the core infrastructure, runtime and operational foundations enabling AI adoption across the organisation.

This role involves collaborating across Cloud Infrastructure, Platform Engineering, SRE, and DevOps to maintain a scalable, resilient and observable AI Platform that supports AI-powered services across the business.

Qualifications

  • 3-8 years of experience in SRE, Platform Engineering, DevOps or Cloud Infra environments.
  • Strong hands-on Terraform and IaC experience.
  • Extensive production Kubernetes experience, preferably EKS.
  • Experience with AWS cloud-native environments.
  • Experience building or maintaining CI/CD pipelines.
  • Experience owning production releases and deployment processes.
  • Strong troubleshooting and incident management experience.
  • Experience with observability and monitoring tools (e.g., Datadog).
  • Linux administration and troubleshooting experience.
  • Python or scripting for automation and tooling.
  • Strong communication and stakeholder management skills.

Responsibilities

  • Build and improve platform reliability, resilience and operational excellence.
  • Own and enhance Terraform-based infrastructure provisioning and management.
  • Manage and support Kubernetes (Amazon EKS) environments in production.
  • Develop and maintain CI/CD pipelines for deployments.
  • Drive observability, monitoring and alerting across the platform.
  • Support incident investigation, root cause analysis and troubleshooting.
  • Coordinate production releases from UAT to Production.
  • Automate operational tasks and tooling.
  • Support AI and agent-based workloads on the platform.

Skills

SRE experience
Platform engineering
DevOps
Cloud infrastructure
Terraform
Kubernetes EKS
AWS
CI/CD pipelines
Release management
Troubleshooting
Observability
Linux administration
Python automation
Communication skills
Stakeholder management
AI/LLM platform experience
Agentic systems experience
Helm
OpenSearch
Databricks
Datadog

Tools

Datadog
Helm
OpenSearch
Databricks

Job description

Overview

We're hiring a Site Reliability Engineer to join a high-performing AI Platform team responsible for building and operating the core infrastructure, runtime and operational foundations that enable AI adoption across the organisation.

This role is ideal for someone who enjoys working across Cloud Infrastructure, Platform Engineering, Site Reliability Engineering, and DevOps disciplines while solving complex operational challenges in a modern cloud-native environment. You will play a key role in ensuring the AI Platform remains scalable, resilient and observable while supporting AI-powered services and applications used across the business.

What You'll Do
  • Build and improve platform reliability, resilience and operational excellence.
  • Own and enhance Terraform-based infrastructure provisioning and management.
  • Manage and support Kubernetes (Amazon EKS) environments in production.
  • Develop and maintain CI/CD pipelines supporting deployment and release activities.
  • Drive observability, monitoring and alerting best practices across the platform.
  • Support incident investigation, root cause analysis and platform troubleshooting.
  • Own and coordinate production releases from UAT through to Production deployment.
  • Automate operational tasks through ing and tooling.
  • Support AI and agent-based workloads running on the platform ecosystem.
Must-Have Skills
  • 3-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps or Cloud Infrastructure environments.
  • Strong hands-on experience with Terraform and Infrastructure as Code.
  • Strong production experience with Kubernetes, ideally Amazon EKS.
  • Strong experience supporting cloud-native environments on AWS.
  • Experience building, maintaining or owning CI/CD pipelines.
  • Experience owning production releases, change processes and deployment activities.
  • Strong troubleshooting and incident management experience.
  • Experience with observability and monitoring tools such as Datadog.
  • Linux administration and troubleshooting experience.
  • Python or ing experience used for automation and operational tooling.
  • Strong communication and stakeholder management skills.
Nice-to-Have Skills
  • AI / LLM platform experience.
  • Agentic systems experience.
  • Helm.
  • OpenSearch.
  • Databricks.
  • AI observability and traceability concepts.
  • Enterprise platform operations.
  • Chaos engineering or resilience testing exposure.

We regret to inform that only shortlisted candidates will be notified .

EA Registration No.: WONG LIN, RACHEL, R25158204

Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA License No. 10C4544

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Application Engineers (AI Platform)
Site Reliability Application Engineers (AI Platform)

ASTEK SINGAPORE INNOVATION TECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 90,000 - 120,000
AI Platform SRE: Cloud, Kubernetes & Observability
AI Platform SRE: Cloud, Kubernetes & Observability

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Platform Engineer
Senior AI Platform Engineer

peoplesearch pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer, Dev Ops
AI Engineer, Dev Ops

Nanyang Technological University Singapore • Singapore

On-site
SGD 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
DevOps / Site Reliability Engineer
DevOps / Site Reliability Engineer

ACCORD INNOVATIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Production Support Analyst (AI Platform)
Production Support Analyst (AI Platform)

TEKsystems Global Services, LLC • Singapore

On-site
SGD 72,000 - 98,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
AI Platform SRE Engineer | Kubernetes & IaC
AI Platform SRE Engineer | Kubernetes & IaC

ASTEK SINGAPORE INNOVATION TECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 90,000 - 120,000
Production Support
Production Support

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 60,000 - 100,000