Platform Engineer

DBiz.ai

Singapore

On-site

SGD 120,000 - 180,000

Full time

2 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

DBiz.ai is seeking a hands-on Platform Operations Engineer to operate and improve Ceph-based storage services for on-prem Kubernetes and OpenShift deployments on a client project going live soon.

You will own day-to-day reliability, perform upgrades and capacity planning, and coordinate hardware maintenance while building runbooks, dashboards, and RCA practices to keep storage robust and scalable.

Qualifications

  • Hands-on Ceph in production environments with day-2 operations
  • Experience with OpenShift Data Foundation (ODF) and storage services for Kubernetes/OpenShift
  • Strong storage operations skills: monitoring, capacity planning, upgrades, expansion, and component replacement
  • Ability to troubleshoot complex storage issues across software, Linux, networks, and hardware
  • Knowledge of storage hardware performance factors: HDD/SSD/NVMe, HBA, firmware
  • Experience with Ceph CSI, RBD, and CephFS for PV management
  • Scripting/automation using Ansible, Python, or shell to reduce manual work
  • Excellent runbooks, documentation, and operational standards

Responsibilities

  • Operate and maintain production Ceph and OpenShift Data Foundation environments
  • Monitor storage health across capacity, latency, throughput, PGs, and device status
  • Perform rebalancing, backfill, scrubbing, recovery actions, and cluster maintenance
  • Plan and execute Ceph/ODF upgrades, patching, expansions, and configuration changes
  • Manage OSD and node replacement with infra teams
  • Diagnose and resolve production incidents (degraded PGs, slow ops, quorum issues)
  • Support Kubernetes/OpenShift storage (Ceph CSI, RBD, CephFS) and PV lifecycle
  • Maintain runbooks, dashboards, alerts, capacity plans, recovery procedures
  • Test recovery from disk/node/service/network failures and improve resilience
  • Participate in incident response, RCA, and drive corrective actions
  • Collaborate with platform/network/infrastructure/application teams for reliability

Skills

Ceph
OpenShift Data Foundation (ODF)
Storage operations
Troubleshooting storage issues
Linux
HDD/SSD/NVMe hardware knowledge
Ceph CSI/RBD/CephFS
Automation (Ansible, Python, shell)
Runbooks and documentation
Incident response/On-call

Job description

We are looking for a hands-on Platform Operations Engineer (Ceph Storage Specialist) to operate and continuously improve Ceph-based storage services that support on-premises Kubernetes and OpenShift platforms for one of our clientele project that is going live soon. This role is operations-focused (production reliability and supportability), covering monitoring, maintenance, upgrades, capacity management, hardware lifecycle, performance troubleshooting, and failure recovery. You will work closely with platform, network, infrastructure, and application teams to keep storage services reliable, scalable, and well-governed.

Key Responsibilities
  • Operate and maintain production Ceph and OpenShift Data Foundation (ODF) environments supporting Kubernetes/OpenShift platforms.
  • Monitor and manage storage health and performance across capacity, latency, throughput, placement groups (PGs), and device status.
  • Perform routine operational tasks including rebalancing, backfill, scrubbing, recovery actions, and general cluster maintenance.
  • Plan and execute Ceph/ODF upgrades, patching, expansions, and configuration changes with minimal service disruption.
  • Manage OSD and node replacement, including coordinating server/disk/firmware maintenance with infrastructure teams.
  • Diagnose and resolve production incidents involving degraded PGs, slow ops, quorum issues, device failures, and network-related storage problems.
  • Support Kubernetes/OpenShift storage consumption patterns, including: (Ceph CSI, RBD (block storage) & CephFS (shared file storage)).
  • Persistent Volume lifecycle troubleshooting (provisioning, attachment, mounting, expansion, and performance)
  • Maintain and improve operational readiness through dashboards, alerts, runbooks, capacity plans, and recovery procedures.
  • Test and validate recovery from disk, node, service, and network failures, and continuously improve resilience.
  • Participate in production incident response, post-incident review, and root-cause analysis (RCA); drive corrective and preventive actions.
  • Collaborate with cross-functional stakeholders (platform, network, infrastructure, application teams) to ensure storage services remain reliable and supportable for clientele platforms.
Required Skills and Experience
  • Hands-on experience operating Ceph in production environments (day-2 operations, troubleshooting, maintenance).
  • Experience operating OpenShift Data Foundation (ODF) and supporting storage services for Kubernetes/OpenShift platforms.
  • Strong experience with storage operations, including monitoring, capacity planning, upgrades, expansion, and component replacement.
  • Proven ability to troubleshoot complex storage issues across software, Linux, networks, and physical hardware.
  • Working knowledge of storage hardware and performance considerations, including HDD, SSD, NVMe, HBA, firmware, and storage-network behaviour.
  • Experience supporting block and shared-file storage using Ceph CSI, RBD, and CephFS, including PV-related troubleshooting.
  • Experience benchmarking storage workloads and analysing performance bottlenecks (latency/throughput, saturation points, noisy neighbour symptoms).
  • Experience in incident handling and operational discipline, including on-call/incident response, triage, and RCA practices.
  • Automation/scripting capability using Ansible, Python, shell, or equivalent tools to reduce manual operations and improve consistency.
  • Strong documentation skills to maintain runbooks, recovery procedures, and operational standards.
Preferred Skills
  • Experience with backup, snapshots, replication, and disaster recovery (DR) operations for Ceph/ODF-backed platforms.
  • Experience building or improving observability for storage platforms (dashboards, alert tuning, actionable SLO/SLA signals).
  • Familiarity with platform/SRE practices such as reliability engineering, change management, and continuous improvement in production environments.
  • Exposure to Kubernetes/OpenShift platform operations beyond storage (helpful for cross-team troubleshooting and incident coordination).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Operations Engineer
Platform Operations Engineer

DIGITAL BIZ SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 110,000 - 140,000
Platform Engineer
Platform Engineer

FPT Asia Pacific • Singapore

On-site
SGD 110,000 - 150,000
Platform Engineer (Ceph Storage Specialist)
Platform Engineer (Ceph Storage Specialist)

XCELLINK PTE. LTD. • Singapore

On-site
SGD 90,000 - 140,000
Platform Engineer (Ceph Storage Specialist)
Platform Engineer (Ceph Storage Specialist)

OX CONSULTANCY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Platform Engineer (Ceph Storage Specialist) - #1678
Platform Engineer (Ceph Storage Specialist) - #1678

JOBSTER PRIVATE LTD. • Singapore

On-site
SGD 90,000 - 130,000
Platform Engineer (Ceph Storage)
Platform Engineer (Ceph Storage)

NSEARCH GLOBAL PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
G05 - Platform Engineer (Ceph Storage Specialist)
G05 - Platform Engineer (Ceph Storage Specialist)

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 130,000
G05 - Platform Engineer (Ceph Storage Specialist)
G05 - Platform Engineer (Ceph Storage Specialist)

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 90,000 - 140,000
Ceph Storage Specialist
Ceph Storage Specialist

FPT ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Ceph Storage Engineer
Ceph Storage Engineer

Optimum Solutions Pte Ltd • Singapore

On-site
SGD 90,000 - 120,000