Senior Database Reliability Engineer

Firmus Technologies Pty Ltd.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Firmus Technologies in Singapore seeks a Senior Database Reliability Engineer to own and evolve the operational datastore powering Firmus AI Cloud and internal platform services. The role sits on bare-metal, self-hosted infrastructure as well as cloud, delivering reliable, scalable database platforms with low cognitive load.

Bring 7+ years in DB reliability, deep PostgreSQL expertise, HA deployments, and automation via Python or Go.

Qualifications

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • 7+ years in database reliability engineering with PostgreSQL in large-scale production.
  • Deep PostgreSQL performance troubleshooting under load (locking, bloat, etc.).
  • Hands-on with PostgreSQL HA/backup stacks (Patroni, pgBackRest, WAL-G).
  • Production Redis or Memcached ownership with HA/clustering.
  • Experience with document/NoSQL stores (MongoDB, Cassandra, DynamoDB).
  • Automation of database work with IaC and Python or Go.

Responsibilities

  • Architecture & standards: define patterns for relational, document, NoSQL, or cache; ensure HA and tenant isolation.
  • Automation & lifecycle: provisioning, upgrades, backups, restore, decommissioning with IaC; run drills.

Skills

PostgreSQL
High concurrency
Multi-tenant
HA/Failover
Backup/Restore
Redis
MongoDB
Python
Go
Infrastructure as Code
Query tuning

Education

Bachelor's degree in computer science or related field

Tools

Patroni
pgBackRest
WAL-G
pgBouncer

Job description

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

ROLE SUMMARY

Firmus Technologies is seeking a Senior Database Reliability Engineer to join our Engineering and Technology team. You will own and evolve the operational datastore platform that Firmus AI Cloud and our internal platform services depend on. The work sits on bare-metal and self-hosted infrastructure as well as cloud, with proven recovery and measurable reliability as first-class product qualities. You build a paved road so product teams can provision and change databases safely, with low cognitive load, instead of waiting on a ticket queue.

KEY RESPONSIBILITIES

Architecture & Standards

Set and improve how we run databases for customer-facing and internal platform workloads. Choose the right pattern for the job: relational, document, NoSQL, or cache. Define connection pooling and database proxies. Define how we scale with vertical growth, read replicas, sharding or partitioning, and multi-site setups. Set HA and failover standards for shared multi-tenant systems. Keep tenant isolation clear and limit blast radius when something fails.

Automation & Lifecycle

Own provisioning, config, upgrades, backup, restore, and decommissioning with infrastructure as code and solid automation. Set RPO and RTO. Prove restore and failover with regular drills. Give product teams safe self-service for routine database work so they do not wait on you for every change.

Reliability & Performance

Own production health across on-premises and cloud database: SLOs, capacity, query performance, and database monitoring. Fix hard production issues such as lock contention, replication lag, memory or storage pressure, slow queries, and storage or network bottlenecks. Join on-call, lead post-mortems, and turn fixes into automation or better runbooks. Track storage and compute cost when you plan capacity.

Security & Compliance

Own access control, encryption in transit and at rest, audit logging, and change control for the databases. Extend Firmus SOC 2 Type 2 and ISO 27001 controls as we add sites and services. Keep backup, restore, and recovery evidence ready for audit. Keep privileged access tight.

Partnership & Escalations

Work with software and platform engineers on schema design, safe migrations, data modelling, and performance reviews before release. Work with data engineering and observability where your databases connect to their systems, such as CDC (Change Data Capture), read replicas, and database telemetry. Run design reviews that raise the bar. Join customer escalations when the issue is database reliability, performance, or recovery.

SKILLS AND EXPERIENCE
  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • 7+ years in database reliability engineering. At least 3 years owning PostgreSQL in large-scale production, such as high concurrency, large datasets, or multi-tenant services. That includes replication, failover, backup and restore, query tuning, and incidents you owned.
  • Deep PostgreSQL performance troubleshooting under load, including planner behaviour, locking, bloat, and connection storms.
  • Hands-on with a PostgreSQL HA and backup stack such as Patroni-style HA, pgBackRest or WAL-G, and pgBouncer or a similar pooler. Has shipped schema or major version changes under load.
  • Production Redis or Memcached ownership. Has run it with HA or clustering, managed memory and eviction under load, and made clear persistence and failover choices.
  • Production experience with at least one document or NoSQL store such as MongoDB, Cassandra, DynamoDB, or similar. Includes HA or replica setup, backup and restore, and at least one failure mode you diagnosed yourself.
  • Automates database work with infrastructure as code and Python or Go. Experience with state
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Database Reliability Engineer
Senior Database Reliability Engineer

Firmus • Singapore

On-site
SGD 120,000 - 180,000
Senior DB Reliability Engineer — PostgreSQL & Cloud
Senior DB Reliability Engineer — PostgreSQL & Cloud

Firmus Technologies Pty Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Database Reliability Engineer — SaaS Infrastructure (SG)
Database Reliability Engineer — SaaS Infrastructure (SG)

Longport Technology HK Limited • Singapore

On-site
SGD 90,000 - 140,000
Senior Database Reliability Engineer: Scale AI Cloud
Senior Database Reliability Engineer: Scale AI Cloud

Firmus • Singapore

On-site
SGD 120,000 - 180,000
Database Administrator — SaaS Infrastructure (SG)
Database Administrator — SaaS Infrastructure (SG)

Longbridge Singapore • Singapore

On-site
SGD 90,000 - 150,000
Senior Database Administrator SQL Server & PostgreSQL
Senior Database Administrator SQL Server & PostgreSQL

IoTalents Pte. Ltd. • Singapore

On-site
SGD 90,000 - 150,000
Principal Engineer, AI Cloud Software
Principal Engineer, AI Cloud Software

Fi • Singapore

On-site
SGD 180,000 - 240,000
Assistant Lead Engineer - Database Developer (Software Development)
Assistant Lead Engineer - Database Developer (Software Development)

Jobline Resources Pte Ltd • Singapore

On-site
SGD 180,000 - 240,000
Database Administrator (Postgres, Observability & DBaaS) - (NPW)
Database Administrator (Postgres, Observability & DBaaS) - (NPW)

Milestone Technologies, Inc. • Singapore

On-site
SGD 180,000 - 240,000
Cloud Senior Database Administrator - AWS
Cloud Senior Database Administrator - AWS

PCCW SOLUTIONS INSYS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000