Site Reliability Engineer

RemotePeople

Singapore

On-site

SGD 120,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

RemotePeople is seeking a senior regional operations engineer to own stability, observability, and automation across AWS regions. You will manage production environments, collaborate with HQ on containerization and middleware governance, and enable independent regional releases with robust incident response.

Responsibilities include day-to-day infrastructure operations, observability platform maintenance, and security controls.

Qualifications

  • Bachelor's degree or higher in CS/SE.
  • 5+ years in operations / SRE / platform engineering.
  • Proficient in multi-region AWS and Kubernetes.
  • Experience with IaC and CI/CD.

Responsibilities

  • Regional infrastructure operations: run day-to-day operations, capacity planning, and cost optimization for the region's AWS resources and Kubernetes clusters.
  • Observability: build and maintain dashboards, alert rules, and log platforms for applications, middleware, and infrastructure.
  • Incident response and drills: regional incident handling; automate mitigation actions and run drills with runbooks.
  • Release and change management: CI/CD pipelines with canary releases and fast rollback; enforce change reviews and IaC governance.
  • Middleware and database operations: backup, recovery, tuning, capacity assessment for regional databases and middleware.
  • Security and access: manage regional IAM accounts; rollout of WAF, gateway, and security capabilities; enforce least privilege and audit trails for data access.

Skills

AWS
Kubernetes
CI/CD
Infra as Code
SRE
Incident response
Observability
Security
Scripting

Education

Bachelor's degree

Tools

Terraform
Jenkins
GitLab CI
Prometheus
Grafana
ELK
MySQL
Redis
Kafka
RocketMQ

Job description

Join an international backend team to own the stability, observability, and automation of the region's production environment. Our client's smart hardware and mobile applications run across multiple AWS regions. You will hold operational access to the regional production environment, handle incidents alongside the regional tech lead, and work with the HQ infrastructure team on containerization, unified gateway, and middleware governance initiatives — giving the region independent release, monitoring, and emergency-response capability.

Responsibilities

  • Regional infrastructure operations: run day-to-day operations, capacity planning, and cost optimization for the region's AWS resources and Kubernetes clusters; keep environments consistent and configuration traceable
  • Observability: build and maintain dashboards, alert rules, and log platforms for applications, middleware (MySQL / Redis / Kafka / RocketMQ), and infrastructure; ensure alerts are accurate and actionable
  • Incident response and drills: take part in regional incident handling; execute and automate mitigation actions (rollback, scaling, degradation switches); run degradation and failure drills and maintain incident runbooks
  • Release and change management: build and maintain CI/CD pipelines supporting canary releases and fast rollback; enforce change review and checklists; advance infrastructure-as-code
  • Middleware and database operations: own backup, recovery, tuning, and capacity assessment for regional databases and middleware; work with developers on slow queries and performance issues
  • Security and access: manage regional IAM accounts and access under company policy; support rollout of WAF, gateway, and other security capabilities; enforce least privilege and audit trails for data access

Requirements

  • Bachelor's degree or above; Computer Science / Software Engineering preferred
  • 5+ years in operations / SRE / platform engineering
  • Proficient with AWS (EC2, ALB, EKS, RDS, ElastiCache, S3, IAM, VPC, CloudWatch), with multi-region operations experience
  • Strong Kubernetes skills: cluster operations, troubleshooting, and resource governance
  • Experience with Terraform or similar infrastructure-as-code tools, and CI/CD tools such as Jenkins / GitLab CI
  • Solid Linux and networking fundamentals; scripting proficiency in at least one of Python / Shell / Go
  • Experience deploying, monitoring, and troubleshooting MySQL, Redis, and Kafka / RocketMQ
  • Familiar with Prometheus / Grafana / ELK or comparable observability stacks
  • Experience handling production incidents and postmortems; able to execute mitigation calmly and methodically under pressure
  • Rigorous and detail-oriented; committed to documented, auditable changes
  • Must be legally authorized to work in the country of employment
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Momcozy • Singapore

On-site
SGD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Momcozy • Singapore

On-site
SGD 120,000 - 180,000
Competitive compensation
Site Reliability Engineer
Site Reliability Engineer

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Regional SRE: Stability, Automation & Observability
Regional SRE: Stability, Automation & Observability

RemotePeople • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

SINGAPORE EXCHANGE LIMITED • Singapore

On-site
SGD 120,000 - 160,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer, Observability ( Contract )
Site Reliability Engineer, Observability ( Contract )

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
G13 - Operations Support Engineer
G13 - Operations Support Engineer

FPT Asia Pacific • Singapore

On-site
SGD 120,000 - 180,000