Application Engineer (Site Reliability)

U3 PROJECTS PTE. LTD.

Singapore

On-site

SGD 90,000 - 150,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

U3 PROJECTS PTE. LTD. is seeking a Site Reliability Engineer to provide SRE and production support for the Ona platform within our Developer platforms.

You will monitor systems, troubleshoot incidents, and perform root-cause analyses to improve reliability and performance. The role requires strong programming skills (JavaScript/Java/Python/.NET), extensive AWS experience, and hands-on work with Docker, Kubernetes, Terraform, and CI/CD tooling.

Qualifications

  • Minimum 5 years of software engineering experience with proficiency in at least one programming language (JavaScript, Java, Python, or .NET).
  • Minimum 2 years of hands-on experience in SRE or production support environment.
  • Minimum 3 years of AWS experience and solid cloud infrastructure knowledge.
  • Minimum 3 years of containerization experience (Docker, Kubernetes, EKS, Helm).

Responsibilities

  • Provide SRE and production support for the Ona platform.
  • Monitor platforms, troubleshoot incidents, perform root-cause analysis and implement reliability improvements.
  • Automate operations, maintain runbooks, and participate in on-call rotations when required.

Skills

JavaScript
Java
Python
.NET
SRE
AWS
Docker
Kubernetes
Terraform
CloudFormation
CI/CD
GitHub Actions
Splunk
CloudWatch
Datadog
Shell scripting
Linux
Incident management

Tools

JFrog

Job description

Summary:
  • The successful candidate will provide Site Reliability Engineering (SRE) and production support for the Ona platform within Developer platforms.
  • Responsibilities include platform monitoring, incident troubleshooting and resolution, root-cause analysis, service reliability improvements, operational automation, runbook maintenance, and on-call support where required.
  • The role will help maintain platform availability, stability, and performance while reducing manual operational effort.
Skillset Requirements.
  • Minimum 5 years of strong software engineering experience with proficiency in at least one programming language, i.e. JavaScript, Java, Python, or .NET.
  • Minimum 2 years of hands‑on experience supporting the reliability and availability of production systems in an SRE or production support environment.
  • Minimum 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
  • Minimum 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous).
  • Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
  • Proficiency with CI/CD workflows and GitHub Actions.
  • Knowledge of artifact repository management systems such as JFrog.
  • Strong Linux administration skills and Shell scripting expertise.
  • Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
  • Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root-cause analysis, and post-incident follow-up.
  • Experience in diagnosing and resolving complex system issues across multiple technology layers.
  • Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
  • Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
  • Willingness to participate in an on-call support rotation for critical production services, where required.
  • Able to optimize developer workflows and enhance developer experience.
  • Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
  • Excellent communication skills to work effectively with diverse engineering teams.
  • Strong team-player mindset, focused on leveraging experience to help the team succeed.
  • Possess positive learning and collaborative mindset.
  • Strong analytical, problem‑solving and troubleshooting skills.
  • Good written and verbal communication skills.
  • Agile, fast learner and able to adapt to changes.

SKILLS: SRE, Production support, Terraform, CloudFormation, AWS, Kubernetes, Docker, CI/CD, Python, Java, JavaScript or.NET, Splunk, CloudWatch, Datadog

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Devops
SRE Devops

Epergne Solutions • Singapore

On-site
SGD 100,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Senior SRE for Cloud Platform & Automation
Senior SRE for Cloud Platform & Automation

U3 PROJECTS PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 130,000
L2 SRE
L2 SRE

NTT Data Singapore • Singapore

On-site
SGD 120,000 - 180,000
SRE Engineer
SRE Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 130,000
We’re Hiring! - Senior SRE Engineer
We’re Hiring! - Senior SRE Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

re-zoo-me • Singapore

On-site
SGD 90,000 - 150,000