Site Reliability Engineer (SRE) / Production Support Engineer

EXPRESS PTE. LTD.

Singapore

On-site

SGD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

EXPRESS PTE. LTD. is seeking an experienced Site Reliability Engineer (SRE) / Production Support Engineer with 10+ years in enterprise infrastructure and 24x7 production support.

The role focuses on high availability, reliability, performance, and security of critical applications with strong hands-on expertise in AIX, Linux, monitoring, automation, and cloud/DevOps. You will collaborate with development, infrastructure, middleware, security, and business teams to identify and resolve production

Qualifications

  • Experience in enterprise infrastructure and 24x7 production support.
  • Strong hands-on in AIX and Linux (RHEL).
  • Proven ability to improve service reliability and incident response.

Responsibilities

  • Define and establish SLIs, SLOs, SLAs and MTTR for enterprise services.
  • Monitor availability, latency, performance, capacity and reliability.
  • Lead incident management and post-incident RCA processes.

Skills

SRE
Production Support
AIX
Linux
Automation
Incident management
Cloud
DevOps
Monitoring
Security

Tools

WebSphere
JBoss
IBM HTTP Server
Apache Tomcat
Grafana
Centreon
AppDynamics
ELK
Nexus/JFrog

Job description

We are seeking an experienced Site Reliability Engineer (SRE) / Production Support Engineer with 10+ years of experience in enterprise infrastructure, production support, system administration, and SRE operations. The ideal candidate will have strong hands-on experience in AIX, Linux/RHEL, application production support, monitoring, automation, incident management, and cloud/DevOps environments. The candidate will be responsible for ensuring high availability, reliability, performance, and security of critical enterprise applications and infrastructure in a 24x7 production support environment. The role requires close collaboration with development, infrastructure, middleware, security, and business teams to identify and resolve production issues and continuously improve service reliability.

Key Responsibilities
Site Reliability Engineering
  • Define and establish SLIs, SLOs, SLAs, error budgets, MTTD, and MTTR for enterprise applications and services.
  • Monitor application availability, latency, performance, capacity, and overall reliability.
  • Identify and implement Golden Signals and observability best practices across applications and infrastructure.
  • Analyze production incidents and recurring issues to identify root causes and implement permanent remediation.
  • Develop automation and engineering solutions to reduce manual operational activities and improve reliability.
  • Participate in 24x7x365 production support and on-call operations.
  • Perform capacity planning and proactively identify infrastructure and application resource requirements.
  • Support highly available and resilient application architectures.
  • Participate in disaster recovery planning, testing, and implementation.
Production & Infrastructure Support
  • Provide L2/L3 production and infrastructure support for critical enterprise applications.
  • Perform administration, troubleshooting, configuration, maintenance, and performance tuning of IBM AIX and RHEL/Linux servers.
  • Support server patching, upgrades, maintenance, and DR activities.
  • Troubleshoot application, middleware, operating system, connectivity, and infrastructure-related issues.
  • Perform middleware administration and restart activities for WebSphere, JBoss, IBM HTTP Server, and Apache Tomcat.
  • Support firewall changes, SSL certificate renewals, SCP configuration, and system-to-system key exchanges.
  • Support application deployment activities across Blue/Green environments.
  • Coordinate application-related changes and deployments across development, infrastructure, middleware, and business teams.
Monitoring & Observability
  • Configure and maintain monitoring and alerting for infrastructure and application services.
  • Work with monitoring and observability platforms such as:
  • Grafana
  • Centreon
  • AppDynamics
  • ELK / Kibana
  • Application and infrastructure logging platforms
  • Analyze system and application metrics, logs, and performance trends.
  • Develop appropriate alerts to proactively identify service degradation and failures.
  • Promote observability practices and help development teams implement effective monitoring.
Incident & Change Management
  • Manage and resolve incidents within defined SLA/OLA timelines.
  • Participate in major incident and emergency response activities.
  • Perform incident investigation, troubleshooting, root-cause analysis, and problem management.
  • Raise and manage Change Requests (CRs) for application deployments, BAU fixes, infrastructure changes, middleware changes, and maintenance activities.
  • Create and manage service requests and incident tickets.
  • Prepare RCA and corrective/preventive action plans for recurring and high-priority incidents.
  • Follow ITIL-based Incident, Change, Problem, and Service Request Management processes.
Automation & DevOps
  • Develop automation scripts using Python and Shell scripting to improve operational efficiency.
  • Automate repetitive infrastructure and production support activities.
  • Work with DevOps tools and technologies including:
  • Git / GitHub / Bitbucket
  • Jenkins
  • Ansible
  • Chef
  • Docker
  • Maven
  • JFrog / Nexus Repository
  • SonarQube
  • Fortify / Nexus IQ
  • Support CI/CD pipelines and application deployment processes.
  • Collaborate with development teams to integrate monitoring, logging, and reliability controls into deployment pipelines.
Cloud & Application Support
  • Provide support for applications hosted on AWS and Pivotal Cloud Foundry (PCF) environments.
  • Work with enterprise application technologies including:
  • IBM WebSphere
  • IBM HTTP Server
  • JBoss
  • Apache Tomcat
  • Java applications
  • Support microservices-based applications and their associated infrastructure.
  • Assist with application migration, optimization, and adoption of new technologies where required.
Security & Compliance
  • Apply system security best practices across AIX and Linux environments.
  • Support security hardening, patching, access management, and vulnerability remediation.
  • Work with security and access management tools such as IBM Tivoli Access Man
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer – Cloud & Automation
Senior Site Reliability Engineer – Cloud & Automation

EXPRESS PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ad Astra Consultants • Singapore

On-site
SGD 90,000 - 130,000
IT Production Support Engineer
IT Production Support Engineer

EVOLUTION RECRUITMENT SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 60,000 - 90,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

re-zoo-me • Singapore

On-site
SGD 90,000 - 150,000
Unix L2/L3 Support Engineer
Unix L2/L3 Support Engineer

AVATAR MODERN TECHNO SERVICES PTE. LTD. • Singapore

On-site
SGD 110,000 - 150,000
Application Support Engineer (Unix/Linux & SQL)
Application Support Engineer (Unix/Linux & SQL)

Luxoft • Singapore

On-site
SGD 78,000 - 112,000
Application Support Engineer (Unix/Linux & SQL)
Application Support Engineer (Unix/Linux & SQL)

Luxoft Singapore • Singapore

On-site
SGD 60,000 - 90,000
Application / Production Support L2 ( with Devops experience)
Application / Production Support L2 ( with Devops experience)

ITCAN PTE. LIMITED • Singapore

On-site
SGD 70,000 - 110,000