SRE Engineer

re-zoo-me

Singapore

Hybrid

SGD 90,000 - 130,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

TAO Digital Solutions in Taipei seeks an experienced SRE engineer to manage multi-cloud operations across AWS, GCP, AliCloud and internal gateways. You will diagnose incidents from intake to RCA, implementing runbooks and blameless postmortems, while coordinating with platform engineering.

The role emphasizes IAM, networking, observability, IaC, and incident discipline, with Singapore anchored APAC coverage.

Qualifications

  • 5+ years in cloud infrastructure support, site reliability, or cloud operations engineering.
  • Deep production expertise in at least one major cloud provider with working competence in a second.
  • Strong IAM fundamentals across providers: role assumption, delegated access and federated identities.
  • Practical Kubernetes operations experience (EKS, GKE, or ACK).
  • Knowledge of routing, DNS, TLS, proxies, CIDR and firewall/ACL models.
  • Scripting in Python or Bash and fluency with provider CLIs.
  • Excellent written English for asynchronous communication.
  • Legally authorized to work in Singapore.

Responsibilities

  • Diagnose incidents across multi-cloud estates from intake to resolution.
  • Write RCA reports and feed learnings into runbooks.
  • Manage IAM roles, trust policies, and access controls across providers.
  • Ensure cost visibility and remediation across environments.
  • Collaborate on platform engineering for capacity and reliability.

Skills

Cloud platforms
Incident management
IAM and access management
Kubernetes operations
Networking fundamentals
Python or Bash scripting
IaC tooling (Terraform/Ansible)
Observability tooling (Prometheus, Spl
Monitoring and RCA

Education

Bachelor’s degree in a relevant field

Tools

Terraform
Ansible
Spinnaker
GCP/AWS/AliCloud CLIs

Job description

Company Description

TAO Digital Solutions is a global technology company headquartered in Silicon Valley, with offices across the US, Canada, Taiwan, India, Australia, New Zealand, and Nigeria. We specialize in product engineering, AI, data, cloud, managed services, and industry solutions across multiple sectors.

Our global team of 4,000+ professionals is passionate about building scalable, innovative, and data-driven solutions that help organizations accelerate digital transformation and unlock business value.

Role Description

We are looking for an SRE engineer to join our team in Taipei. In this role, you will diagnose and resolve issues across the multicloud estate: AWS, GCP, AliCloud, and the internal cloud gateway and control planes. You will take incidents from page or ticket intake through to resolution and written RCA, escalating to platform engineering only when a confirmed defect or capacity change is required.

Singapore is the anchor site for this role, covering the APAC estate at its point of origin.

Cloud platform core — must have
  • Account lifecycle across providers: onboarding, org hierarchy placement, service enablement, decommission
  • Provisioning failure diagnosis and stuck-state recovery, from request intake to delivered resource
  • Quota management: request, validation, pre-requisite checks, and capacity escalation to the provider
  • IAM and access: roles, trust policies, federated and short-lived credentials, SSO role mapping
  • Cost visibility and remediation: identifying waste, right-sizing, and tracking realized savings
Multi-cloud — must have at least two
  • AWS: EC2, VPC, NLB and ALB, Route 53, subnet and address-space reconciliation, IAM roles anywhere
  • GCP: GCE, service accounts, vulnerability remediation workflows, project and folder structure
  • AliCloud: region-specific service availability and version constraints, STS and RAM
  • Cross-provider differences in quota, IAM, and network models. Knowing where the three diverge is the point of the role
Platform tooling — must have
  • Infrastructure as code: Terraform or Ansible, including drift detection and blast-radius management
  • Deployment and release tooling such as Spinnaker
  • Compass compliance onboarding, backup services, image and golden-AMI lifecycle
  • Internal cloud inventory and query tooling for account, resource, and spend reporting
Observability — must have
  • Prometheus and Grafana: dashboard authorship and alert definition that is actionable, not noise
  • Splunk: forwarding, query authorship, and use in live incident diagnosis
  • Defining SLIs for provisioning success, gateway availability, and request latency
Networking — must have
  • Cross-environment DNS resolution and connectivity to managed database services
  • Load balancer reconcile loops and subnet annotation behavior
  • Gateway and proxy endpoint troubleshooting, including destination allow-listing and egress paths
  • Outage triage and formal RCA authorship
Automation and code — must have
  • Python or Go sufficient to ship tooling that goes through peer review, not throwaway scripts
  • Converting repeat manual remediation into runbook automation or self-service paths
  • Git-based operational workflow, GitHub Actions, PagerDuty
Incident discipline — must have
  • Run an incident end to end: triage, mitigate, verify recovery against the customer-facing signal, resolve
  • Journal every investigation in writing as you work, so the next shift continues instead of restarting
  • Blameless RCA authorship and promotion of findings into the shared runbook library
  • Clear written English. Most hand-offs happen in writing across time zones
Required Qualifications
  • 5+ years in cloud infrastructure support, site reliability, or cloud operations engineering. The multi-provider scope sets this bar above a single-cloud equivalent role.
  • Deep production expertise in at least one major cloud provider, with demonstrated working competence in a second. Depth in AWS or GCP is the most common qualifying profile.
  • Strong identity and access management fundamentals that transfer across providers: role assumption and delegation, service and machine identities, policy evaluation and troubleshooting, and federated or directory-sourced group mapping.
  • Practical Kubernetes operations experience (EKS, GKE, or ACK) — inspecting pod and node state, diagnosing volume and networking failures, and interpreting cluster events.
  • Working knowledge of enterprise network fundamentals: routing, DNS, TLS, proxies, CIDR planning, and firewall/ACL models.
  • Scripting proficiency in Python or Bash, plus fluency with provider CLIs.
  • Excellent written English — the majority of support happens asynchronously in text, and clarity directly determines resolution speed.
  • Legally authorized to work in Singapore.
Preferred Qualifications
  • Certification in one or more of: AWS Solutions Architect (Associate/Professional), Google Professional Cloud Architect, Alibaba Cloud ACP.
  • Prior AliCloud experience, or demonstrated rapid ramp on an unfamiliar cloud provider. AliCloud talent is scarce in the market; evidence of learning a new provider quickly is an acceptable and expected substitute.
  • Experience supporting internal developer platforms in a large enterprise.
  • Familiarity with hybrid connectivity between corporate networks and public cloud.
  • Infrastructure-as-code exposure (Terraform, CloudFormation).
  • Observability tooling experience (Splunk, Datadog, CloudWatch, Cloud Logging).
  • Prior follow-the-sun or multi-region support rotation experience.
  • Familiarity with China-region cloud operations and their regulatory constraints.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineer - Multi-Cloud Reliability & Incident Response
SRE Engineer - Multi-Cloud Reliability & Incident Response

re-zoo-me • Singapore

Hybrid
SGD 90,000 - 130,000
Cloud Operations Support Engineer
Cloud Operations Support Engineer

SEDHA CONSULTING PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Momcozy • Singapore

On-site
SGD 120,000 - 180,000
Competitive compensation
[LPS] Service Delivery Lead - Cloud
[LPS] Service Delivery Lead - Cloud

LPS • Singapore

On-site
SGD 75,000 - 100,000
Cloud Engineer
Cloud Engineer

WORLD PARTNERS SOLUTION PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior Software Engineer (Operations)
Senior Software Engineer (Operations)

UFINITY PTE LTD • Singapore

On-site
SGD 120,000 - 180,000
Senior Engineer
Senior Engineer

Starhub Ltd • Singapore

On-site
SGD 90,000 - 130,000
Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
Technical Manager
Technical Manager

COMBUILDER PTE LTD • Singapore

On-site
SGD 120,000 - 190,000
Attractive Job Opening for Senior L1/L2 Support Engineer in Singapore
Attractive Job Opening for Senior L1/L2 Support Engineer in Singapore

TALENT XPERTS PTE. LTD. • Singapore

On-site
SGD 67,000 - 100,000