Application Engineer (Site Reliability)

U3 INFOTECH PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

U3 INFOTECH PTE. LTD. seeks a Site Reliability Engineer to provide SRE and production support for the Ona platform within Developer platforms. You will monitor the platform, troubleshoot incidents, perform root-cause analysis, and drive reliability improvements.

You will automate operations, maintain runbooks, participate in on-call rotations, and collaborate with diverse engineering teams to optimize availability, stability and performance while reducing manual effort.

Qualifications

  • Minimum 5 years of strong software engineering experience with proficiency in at least one programming language.
  • Minimum 2 years of hands‑on experience supporting the reliability and availability of production systems in an SRE or production support environment.
  • Minimum 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
  • Minimum 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous).
  • Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
  • Proficiency with CI/CD workflows and GitHub Actions.
  • Knowledge of artifact repository management systems such as JFrog.
  • Strong Linux administration skills and Shell scripting expertise.
  • Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
  • Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root‑cause analysis, and post‑incident follow‑up.
  • Experience in diagnosing and resolving complex system issues across multiple technology layers.
  • Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
  • Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
  • Willingness to participate in an on‑call support rotation for critical production services, where required.
  • Able to optimize developer workflows and enhance developer experience.
  • Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
  • Excellent communication skills to work effectively with diverse engineering teams.
  • Strong team‑player mindset, focused on leveraging experience to help the team succeed.
  • Possess positive learning and collaborative mindset.
  • Strong analytical, problem‑solving and troubleshooting skills.
  • Good written and verbal communication skills.
  • Agile, fast learner and able to adapt to changes.

Responsibilities

  • Provide SRE and production support for the Ona platform within Developer platforms.
  • Monitor the platform and troubleshoot incidents, perform root-cause analysis, and drive reliability improvements.
  • Automate operational tasks, maintain runbooks, and participate in on-call rotations as needed.
  • Ensure platform availability, stability and performance while reducing manual effort.

Skills

SRE mindset
Strong communication
Team collaboration
Agile learning
Problem solving

Tools

AWS
Kubernetes
Docker
Terraform
CloudFormation
CI/CD pipelines
GitHub Actions
JFrog
Linux
Shell scripting
CloudWatch
Splunk
Datadog
EKS
Helm

Job description

Summary:

  • The successful candidate will provide Site Reliability Engineering (SRE) and production support for the Ona platform within Developer platforms.
  • Responsibilities include platform monitoring, incident troubleshooting and resolution, root-cause analysis, service reliability improvements, operational automation, runbook maintenance, and on-call support where required.
  • The role will help maintain platform availability, stability, and performance while reducing manual operational effort.

Skillset Requirements

  • Minimum 5 years of strong software engineering experience with proficiency in at least one programming language, i.e. JavaScript, Java, Python, or .NET.
  • Minimum 2 years of hands‑on experience supporting the reliability and availability of production systems in an SRE or production support environment.
  • Minimum 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
  • Minimum 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous).
  • Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
  • Proficiency with CI/CD workflows and GitHub Actions.
  • Knowledge of artifact repository management systems such as JFrog.
  • Strong Linux administration skills and Shell scripting expertise.
  • Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
  • Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root‑cause analysis, and post‑incident follow‑up.
  • Experience in diagnosing and resolving complex system issues across multiple technology layers.
  • Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
  • Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
  • Willingness to participate in an on‑call support rotation for critical production services, where required.
  • Able to optimize developer workflows and enhance developer experience.
  • Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
  • Excellent communication skills to work effectively with diverse engineering teams.
  • Strong team‑player mindset, focused on leveraging experience to help the team succeed.
  • Possess positive learning and collaborative mindset.
  • Strong analytical, problem‑solving and troubleshooting skills.
  • Good written and verbal communication skills.
  • Agile, fast learner and able to adapt to changes.

SKILLS: SRE, Production support, Terraform, CloudFormation, AWS, Kubernetes, Docker, CI/CD, Python, Java, JavaScript or.NET, Splunk, CloudWatch, Datadog
U3 Privacy Notice:
Please refer to U3’s Privacy Notice for Job Applicants/Seekers at https://u3infotech.com/privacy-notice-job-applicants/.When you apply, you voluntarily consent to the collection, use and disclosure of your personal data for recruitment/employment and related purposes.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Application Engineer (Site Reliability)
Application Engineer (Site Reliability)

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Senior SRE for Cloud Platform & Automation
Senior SRE for Cloud Platform & Automation

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Platform SRE Engineer – Reliability & Automation
Platform SRE Engineer – Reliability & Automation

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Application Engineer
Application Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 60,000 - 85,000
Site Reliability Engineer, Observability ( Contract )
Site Reliability Engineer, Observability ( Contract )

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
SRE Devops
SRE Devops

Epergne Solutions • Singapore

On-site
SGD 100,000 - 180,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
SRE Engineer
SRE Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
DevOps / Site Reliability Engineer
DevOps / Site Reliability Engineer

ACCORD INNOVATIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000