Site Reliability Engineer

Applicantz

Singapore

On-site

SGD 120,000 - 180,000

Full time

3 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Applicantz is seeking a Site Reliability Engineer (SRE) to ensure reliability, availability, and performance of our systems. The role blends software engineering with operations, leveraging automation, monitoring, and data-driven analysis to improve reliability while maintaining development velocity.

In this structure, SREs act as reliability owners and domain practitioners, supporting platform and product teams with SRE and DevOps responsibilities, guided by senior leadership to align standards

Qualifications

  • Strong knowledge of AWS core services.
  • Experience with multi-AZ and multi-region architectures.
  • Proficient networking in AWS (VPC, subnets, routing, security groups).
  • Experience with well-architected framework pillars.
  • Proven incident response and postmortem leadership.
  • Hands-on with IaC and monitoring tools.
  • Own end-to-end reliability with SLAs/SLOs.

Responsibilities

  • Ensure reliability, availability, and performance of systems.
  • Apply software engineering to operations with automation.
  • Lead major incident response and blameless postmortems.
  • Standardize observability across environments.

Skills

AWS core services
Multi-AZ/region design
Networking in AWS
Well-Architected framework
Incident management
Observability
Automation & SRE practices

Tools

Terraform
CloudFormation
CDK
CloudWatch
Prometheus
Grafana
Datadog
Dynatrace
OpenTelemetry

Job description

  • Advance knowledge of core AWS services: EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM.
  • Designing multi-AZ and multi-region highly available architectures.
  • Strong understanding of networking in AWS (subnets, routing tables, NAT, security groups, NACLs, VPC peering, PrivateLink).
  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK
  • Hands-on experience with CloudWatch, Prometheus, Grafana, Datadog, Dynatrace, or OpenTelemetry
  • Modular IaC design patterns and state management best practices.
  • Own end-to-end system reliability, availability, and performance using clearly defined SLAs, SLOs, and SLIs, with continuous monitoring and proactive improvement of service health.
  • Establish and govern error budget policies in partnership with engineering leadership to balance release velocity with reliability, using error budgets to inform prioritization and release readiness decisions.
  • Lead major and complex incident response efforts, collaborate during customer-impacting events, and drive blameless postmortems to ensure systemic corrective actions are implemented with urgency.
  • Standardize and enhance observability across environments through robust monitoring, logging, and tracing frameworks using tools such as Dynatrace, CloudWatch, and OpenTelemetry.
Role Summary

The Site Reliability Engineer (SRE) ensures the reliability, availability, and performance of systems and platform services through a balance of engineering and operational excellence. The SRE applies software engineering principles to operations, using automation, monitoring, and data-driven analysis to improve reliability while enabling development velocity.

In the current structure, the SREs operate as both reliability owners and domain practitioners, supporting platform and product engineering teams across SRE and DevOps responsibilities. They are guided by a Senior Principal SRE, who provides organizational alignment, establishes common standards, and ensures consistency across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SEVEN HILLS CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer — Reliability, Automation & Observability
Site Reliability Engineer — Reliability, Automation & Observability

Applicantz • Singapore

On-site
SGD 120,000 - 180,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer - HM: Mukesh
Site Reliability Engineer - HM: Mukesh

NTT Data Singapore • Singapore

On-site
SGD 80,000 - 120,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ad Astra Consultants • Singapore

On-site
SGD 90,000 - 130,000
SRE Lead: Scale, Reliability & Observability on AWS
SRE Lead: Scale, Reliability & Observability on AWS

kidentify pte. ltd. • Singapore

On-site
SGD 80,000 - 120,000
Senior Platform Engineer / Site Reliability Engineer (SRE)
Senior Platform Engineer / Site Reliability Engineer (SRE)

AMBITION GROUP SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Cloud-Scale SRE: Reliability, Observability & Incident Mastery
Cloud-Scale SRE: Reliability, Observability & Incident Mastery

SEVEN HILLS CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000