Senior Site Reliability Engineer

Kidentify

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kidentify seeks a Senior Site Reliability Engineer to own and improve the reliability of our platform. You will manage infrastructure on AWS and Kubernetes, build strong observability, and lead incident response for a globally scaled service.

You will collaborate with engineering to bake reliability into design, automate toil away, and optimize deployment safety with robust CI/CD and environment controls. This is a hands-on role requiring coding and systems thinking.

Qualifications

  • 5+ years of production ownership in infrastructure, platforms, SRE, or software with ownership responsibilities.
  • Strong experience running production systems on AWS.
  • Hands-on experience with Kubernetes and container workloads.
  • Experience with infrastructure as code, preferably Terraform.
  • Experience designing observability stacks with Prometheus, Grafana, OpenTelemetry, or equivalents.
  • Strong understanding of distributed systems, failure modes, and production debugging.
  • Experience building or improving CI/CD systems and release workflows.
  • Ability to write code and automation in Go, Python, or TypeScript.
  • Good judgment during incidents with a practical mindset on tradeoffs and recovery.

Responsibilities

  • Own the reliability, availability, and performance of the systems behind k-ID's platform and public APIs.
  • Design and improve scalable infrastructure on AWS and Kubernetes to support growth and global workloads.
  • Build and maintain observability across logs, metrics, tracing, alerting, and service health.
  • Improve deployment safety through CI/CD workflows, release controls, and environment consistency.
  • Lead incident response and production readiness practices, including runbooks, on-call hygiene, and postmortems.
  • Reduce operational toil by automating repetitive tasks and improving internal tooling.
  • Partner with engineering teams to embed reliability and operability from the start of service design.
  • Strengthen platform security and hygiene across access controls, secrets handling, and hardening.
  • Continuously improve system performance and cost awareness without sacrificing reliability.

Skills

AWS
Kubernetes
CI/CD
Go
Python
TypeScript
Observability
Distributed systems
Incident response

Tools

Terraform
Prometheus
Grafana
OpenTelemetry

Job description

About the role

We are hiring a Senior Site Reliability Engineer to help make k-ID reliable at scale.

This role sits in the middle of our production backbone. You will own and improve the systems that keep our platform available, observable, secure, and resilient as traffic grows and our client base expands globally. You will work across infrastructure, tooling, deployment workflows, incident response, and systems design to make sure we can scale without breaking.

This is not a ticket closing operations role. We want someone who can look at a system, find the weak points, and harden it. Someone who cares about failure modes, blast radius, deployment safety, recovery time, cost discipline, and the realities of running production systems under pressure. You should be comfortable writing code, automating away toil, and partnering closely with engineers to improve reliability through better architecture and better operating practices.

Responsibilities
  • Own the reliability, availability, and performance of the systems behind k-ID's platform and public APIs
  • Design and improve scalable infrastructure on AWS and Kubernetes that can support high growth, uneven traffic, and global production workloads
  • Build and maintain strong observability across logs, metrics, tracing, alerting, and service health so issues are caught early and investigated quickly
  • Improve deployment safety through better CI and CD workflows, release controls, rollback paths, and environment consistency
  • Drive incident response and production readiness practices, including runbooks, on call hygiene, postmortems, capacity planning, and resilience testing
  • Reduce operational toil by automating repetitive work and improving internal tooling for developers and operators
  • Partner with engineering teams to embed reliability and operability into service design from the start, not after something fails in production
  • Strengthen platform security and infrastructure hygiene across access controls, secrets handling, system hardening, and production safeguards
  • Continuously improve system performance, resource efficiency, and cost awareness without compromising reliability
Qualifications
  • 5+ years of experience in infrastructure, platform engineering, site reliability engineering, or software engineering with meaningful production ownership
  • Strong experience running production systems in AWS
  • Strong hands on experience with Kubernetes and container based workloads
  • Experience with infrastructure as code, preferably Terraform
  • Experience designing and operating observability stacks using tools such as Prometheus, Alertmanager, Grafana, OpenTelemetry, or equivalent systems
  • Strong understanding of distributed systems, failure modes, service reliability, and production debugging
  • Experience building or improving CI and CD systems and release workflows in modern engineering environments
  • Ability to write code and automation in one or more languages such as Go, Python, or TypeScript
  • Good judgment during incidents and a practical mindset around tradeoffs, risk, and recovery
  • Clear written and verbal communication skills with the ability to work effectively in a remote team
  • Startup experience is a plus, especially in environments where systems and processes are still being built
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

kidentify pte. ltd. • Singapore

On-site
SGD 80,000 - 120,000
Head of Site Reliability Engineering (SRE) & Information Security
Head of Site Reliability Engineering (SRE) & Information Security

Kristal Advisors (Sg) Pte. Ltd. • Singapore

On-site
SGD 180,000 - 260,000
Senior Site Reliability Engineer - Scale & Resilience Leader
Senior Site Reliability Engineer - Scale & Resilience Leader

Kidentify • Singapore

On-site
SGD 120,000 - 180,000
Lead Site Reliability Engineer - SG
Lead Site Reliability Engineer - SG

k-ID • Singapore

On-site
SGD 100,000 - 150,000
Employee Stock Ownership Plan
Comprehensive health coverage
Mental Health support
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA Public Security • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
DevOps Lead (Platform & Infrastructure Engineering)
DevOps Lead (Platform & Infrastructure Engineering)

CDG ZIG PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
SRE Lead: Scale, Reliability & Observability on AWS
SRE Lead: Scale, Reliability & Observability on AWS

kidentify pte. ltd. • Singapore

On-site
SGD 80,000 - 120,000