SRE

Tech Aalto Pte ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Tech Aalto Pte Ltd seeks Site Reliability Engineers to treat operations as a software problem and build automation, tooling, and agentic workflows to keep production healthy and boring to run.

Role requires strong SRE experience, code-first mindset, and hands-on with Kubernetes, IaC, CI/CD, and modern cloud environments in regulated settings.

Qualifications

  • 5+ years in SRE, platform, or infrastructure engineering with a track record of replacing manual work with code.
  • Strong programming ability in at least one modern language (Go, Python, Kotlin, TypeScript, Rust).
  • AI-native ways of working: orchestrating agents for ops workflows, not just AI-assisted autocomplete.
  • Deep hands-on with Kubernetes, IaC (Terraform or equivalent), CI/CD, and observability stacks.

Responsibilities

  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn incidents into durable fixes in code or automation.
  • Build the platform and tooling that make services easy to deploy, observe, and operate: CI/CD, infrastructure-as-code, observability stacks, runbooks-as-code.
  • Apply AI agentically across operations (triage, root-cause analysis, remediation, change review) and contribute to internal agentic ecosystem.
  • Design and integrate the systems underneath our services: messaging (Kafka), orchestration (Kubernetes), and performance-sensitive infrastructure.
  • Partner with product engineers on release readiness, rollout strategy, and production hardening before things ship.
  • Continuously reduce toil: measure it, attack it with code, and raise the floor on what 'easy to maintain' looks like.

Skills

SRE experience
Go
Python
TypeScript
Kubernetes
Terraform
CI/CD
Observability
Incident response

Tools

Kafka

Job description

J ob Summary

We are hiring Site Reliability Engineers who treat operations as a software problem. You'll keep production healthy, but more importantly you'll build the automation, tooling, and agentic workflows that make running our systems boring and predictable. This is an engineering role - if your instinct on a recurring issue is to write code that removes it, you'll fit in well.

Our client operates in a regulated capital-markets environment, so the bar for reliability, security, and operational rigour is high.

Job Responsibilities
  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn every incident into a durable fix in code or automation.
  • Build the platform and tooling that make services easy to deploy, observe, and operate: CI/CD, infrastructure-as-code, observability stacks, runbooks-as-code.
  • Apply AI agentically across operations (triage, root-cause analysis, remediation, change review) and contribute to our internal agentic ecosystem.
  • Design and integrate the systems underneath our services: messaging (e.g. Kafka), orchestration (e.g. Kubernetes), and performance-sensitive infrastructure.
  • Partner with product engineers on release readiness, rollout strategy, and production hardening before things ship.
  • Continuously reduce toil: measure it, attack it with code, and raise the floor on what "easy to maintain" looks like.
Job Requirements
  • 5+ years in SRE, platform, or infrastructure engineering, with a clear track record of replacing manual work with code
  • Strong programming ability in at least one modern language (e.g. Go, Python, Kotlin, TypeScript, Rust, etc), you write production code, not just glue scripts
  • AI-native ways of working: real experience orchestrating agents for ops workflows, not just using AI for autocomplete
  • Deep hands-on with Kubernetes, IaC (Terraform or equivalent), CI/CD, and modern observability (metrics, logs, traces)
  • Production experience on a major cloud: GCP preferred, AWS acceptable
  • Solid foundations in distributed systems and the failure modes that matter in production
  • Incident-response maturity: calm under pressure, sharp on root cause, disciplined about follow-through
  • Comfort in complex, regulated environments
Nice to Have
  • Familiarity with the FIX protocol or capital-markets domain
  • Experience building internal developer platforms or self-service tooling consumed by other engineers

Confidentiality is assured, and only shortlisted candidates will be notified for interviews.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Kubernetes & Site Reliability Engineer (SRE)
Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ad Astra Consultants • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
SRE Engineer
SRE Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Singapore Exchange Limited • Singapore

Hybrid
SGD 180,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

SEVEN HILLS CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Kubernetes & Site Reliability Engineer (SRE)
Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology
SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology

DBS Bank • Singapore

On-site
SGD 300,000 - 520,000