Site Reliability Engineer (SRE)

ASTEK SINGAPORE INNOVATION TECHNOLOGY PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Astek Singapore Innovation Technology Pte. Ltd. is seeking an experienced Production Support / Site Reliability Engineer (SRE) to ensure mission-critical systems stay reliable in cloud and containerised environments.

You will handle incident management, root cause analysis, and stakeholder support, while developing automation and monitoring to improve service reliability. A strong Linux/Unix, cloud, and Kubernetes background is required.

Qualifications

  • Degree in Computer Science or Information Technology or related field.
  • 5+ years in Production/Application Support or SRE roles.
  • Hands-on with business-facing applications and users.

Responsibilities

  • Provide day-to-day production and application support for mission-critical systems.
  • Investigate and resolve complex issues across multiple technology layers.
  • Manage incident triage, incident management, problem management, and RCA.
  • Monitor health, performance, and workloads to ensure service reliability.
  • Troubleshoot across Linux/Unix, cloud, containerised, and app environments.
  • Develop and maintain Bash/Shell scripts for automation.
  • Collaborate with business users, engineering, and stakeholders to minimise disruption.
  • Identify opportunities to improve reliability, monitoring, automation, and processes.

Skills

Production Support
Site Reliability Engineering
Incident management
Troubleshooting
Automation
Shell scripting
Stakeholder management
Agentic AI

Education

Bachelor's degree in Computer Science or Information Technology

Tools

Control-M
Unix/Linux
Bash/Shell scripting
AWS
Azure
Kubernetes
Datadog
Terraform
Agentic AI

Job description

Production Support / Site Reliability Engineer (SRE)
Role Overview

We are looking for an experienced Production Support / Site Reliability Engineer (SRE) with experience in Agentic AI to support mission-critical, business-facing applications. This is a hands-on, techno-functional role covering production operations, incident resolution, system reliability, and stakeholder support across modern cloud and containerised environments.

Key Responsibilities
  • Provide day-to-day production and application support for mission-critical, business-facing systems.
  • Investigate and resolve complex application and infrastructure issues across multiple technology layers.
  • Manage incident triage, incident management, problem management, and root cause analysis.
  • Monitor application health, system performance, batch processes, and scheduled workloads to maintain service reliability.
  • Troubleshoot issues across Linux/Unix, cloud, containerised, and application environments.
  • Develop and maintain Bash/Shell scripts to support operational activities and automation.
  • Work closely with business users, engineering teams, and other stakeholders to resolve production issues and minimise service disruption.
  • Identify opportunities to improve system reliability, monitoring, automation, and operational processes.
Requirements
  • Degree in Computer Science, Information Technology, or a related discipline.
  • At least 5 years of experience in Production/Application Support or Site Reliability Engineering (SRE).
  • Strong hands-on experience supporting business-facing applications and users.
  • Proficiency in Control-M, Unix/Linux, Bash, and Shell scripting.
  • Experience with AWS and/or Azure and cloud-native environments.
  • Hands-on experience with Kubernetes and containerised applications.
  • Familiarity with operational and infrastructure tools such as AutoSys, Datadog, and Terraform.
  • Strong troubleshooting skills across application and infrastructure layers.
  • Experience with incident and problem management.
  • Strong analytical, communication, and stakeholder management skills.
  • Proactive, collaborative, and adaptable approach to working in fast-paced environments
  • Exposure to Agentic AI technologies and AI-enabled operational use cases.
Key Technologies

Control-M | Unix/Linux | Bash/Shell | AWS | Azure | Kubernetes | AutoSys | Datadog | Terraform | Cloud-Native | Agentic AI

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Astek • Singapore

On-site
SGD 110,000 - 150,000
Application Engineer (Site Reliability)
Application Engineer (Site Reliability)

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 130,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Production Support Engineer
Production Support Engineer

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 90,000 - 150,000
Production Support
Production Support

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 60,000 - 90,000
Kubernetes & Site Reliability Engineer (SRE)
Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
Senior SRE & Production Reliability Engineer (Agentic AI)
Senior SRE & Production Reliability Engineer (Agentic AI)

Astek • Singapore

On-site
SGD 110,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

XCELLINK PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000