Site Reliability Engineer

SEVEN HILLS CONSULTING PTE. LTD.

Singapore

On-site

SGD 90,000 - 130,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Seven Hills Consulting Pte. Ltd. seeks an experienced Site Reliability Engineer to ensure the reliability, availability and performance of core systems and platform services.

The role blends software engineering with operations, emphasizing automation, monitoring and data‑driven analysis to improve reliability while preserving velocity. As reliability owners and domain practitioners, SREs support platform and product teams, guided by senior leadership to establish standards, observability,

Qualifications

  • Experience with AWS core services (EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM).
  • Designing multi-AZ and multi-region architectures for high availability.
  • Networking in AWS (subnets, routing, NAT, security groups, NACLs, VPC peering, PrivateLink).
  • Experience with well-architected framework pillars (reliability, security, cost optimization).
  • Designing fault-tolerant, horizontally scalable systems.
  • Proficiency with Terraform, CloudFormation, or CDK.
  • Hands-on with CloudWatch, Prometheus, Grafana, Datadog, Dynatrace, OpenTelemetry.
  • Modular IaC design patterns and state management best practices.

Responsibilities

  • Own end-to-end system reliability, availability, and performance with defined SLAs/SLOs/SLIs and proactive health improvements.
  • Establish and govern error budget policies to balance release velocity and reliability.
  • Lead major incident response efforts and drive blameless postmortems for systemic improvements.
  • Standardize and enhance observability across environments using monitoring and tracing tools.

Job description

Technical Skills
  • Advance knowledge of core AWS services: EC2, ECS/EKS, Lambda, S3, RDS/Aurora, DynamoDB, VPC, ELB/ALB/NLB, Route53, IAM.
  • Designing multi-AZ and multi-region highly available architectures.
  • Broad understanding of networking in AWS (subnets, routing tables, NAT, security groups, NACLs, VPC peering, PrivateLink).
  • Experience with well-architected framework pillars (especially reliability, security, cost optimization).
  • Designing fault-tolerant and horizontally scalable systems
  • Advanced proficiency in Terraform, CloudFormation, or CDK
  • Hands-on experience with CloudWatch, Prometheus, Grafana, Datadog, Dynatrace, or OpenTelemetry
  • Modular IaC design patterns and state management best practices.
  • Own end-to-end system reliability, availability, and performance using clearly defined SLAs, SLOs, and SLIs, with continuous monitoring and proactive improvement of service health.
  • Establish and govern error budget policies in partnership with engineering leadership to balance release velocity with reliability, using error budgets to inform prioritization and release readiness decisions.
  • Lead major and complex incident response efforts, collaborate during customer-impacting events, and drive blameless postmortems to ensure systemic corrective actions are implemented with urgency.
  • Standardize and enhance observability across environments through robust monitoring, logging, and tracing frameworks using tools such as Dynatrace, CloudWatch, and OpenTelemetry.
Role Summary

The Site Reliability Engineer (SRE) ensures the reliability, availability, and performance of systems and platform services through a balance of engineering and operational excellence. The SRE applies software engineering principles to operations, using automation, monitoring, and data-driven analysis to improve reliability while enabling development velocity.

In the current structure, the SREs operate as both reliability owners and domain practitioners, supporting platform and product engineering teams across SRE and DevOps responsibilities. They are guided by a Senior Principal SRE, who provides organizational alignment, establishes common standards, and ensures consistency across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - HM: Mukesh
Site Reliability Engineer - HM: Mukesh

NTT Data Singapore • Singapore

On-site
SGD 80,000 - 120,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

PURVIEW ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Purview Asia Pacific • Singapore

On-site
SGD 120,000 - 180,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA Public Security • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
Site Reliability Engineer(Senior SRE)
Site Reliability Engineer(Senior SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 180,000 - 300,000