Cloud Storage Engineer

Bitdeer Group

Singapore

On-site

SGD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer is seeking an experienced Storage Infrastructure Engineer to design, implement, and optimize distributed and parallel file systems for its GPU cloud. You will drive end-to-end storage integration, ensuring multi-tenant isolation, provisioning, and lifecycle management across a multi-region setup.

Responsibilities include tuning throughput/latency for large-scale AI workloads, managing golden images and CUDA drivers, and building monitoring and runbooks.

Qualifications

  • 3+ years (Senior 6+) in storage engineering or platform infrastructure.
  • Strong understanding of distributed file system internals — data / metadata separation, replication, consistency models.
  • Proven experience integrating and operating distributed storage in production (Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS, MinIO).
  • Performance tuning for high-throughput / parallel I/O; familiarity with NVMe, RDMA / RoCE storage networking, and caching is a strong plus.
  • Strong Linux systems depth and automation skills (Python / Go, CI / CD).
  • HPC / AI storage or multi-region storage experience a strong plus

Responsibilities

  • Design and integrate distributed / parallel file systems into the GPU cloud, optimized for AI training / inference I/O patterns.
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform / control plane.
  • Tune storage throughput and latency for large-scale parallel access (dataset loading, checkpointing); benchmark across GPU SKUs and workloads.
  • Architect multi-region storage: data locality, replication / consistency, durability (failure domains), and cross-region access.
  • Own golden images, templates, GPU drivers / CUDA, and the container / image registry, including versioned release and multi-region distribution. (Secondary scope.)
  • Build monitoring, capacity planning, and runbooks; eliminate single points of failure.
  • Partner with Compute (delivery), Network (storage fabric / RDMA), and Control Plane (provisioning / quota) teams.

Skills

Distributed FS
Linux
Python
Go
CI/CD
NVMe
RDMA
RoCE
GPU drivers
CUDA
Container registry
Automation
Performance tuning
Multi-region storage

Tools

Ceph
Lustre
GPFS / Spectrum Scale
BeeGFS
JuiceFS
MinIO

Job description

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About the team:

GPU training and inference at 10,000+ GPU, multi-region scale depend on high-throughput, low-latency storage that can sustain massive parallel I/O. We are looking for an engineer who deeply understands distributed file systems and can integrate distributed / parallel storage systems into our GPU cloud — covering performance, multi-tenancy, and reliability — while also owning the image / driver / registry pipeline on the node-delivery critical path.

What you will be responsible for:
  • Design and integrate distributed / parallel file systems (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS) into the GPU cloud, optimized for AI training / inference I/O patterns.
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform / control plane.
  • Tune storage throughput and latency for large-scale parallel access (dataset loading, checkpointing); benchmark across GPU SKUs and workloads.
  • Architect multi-region storage: data locality, replication / consistency, durability (failure domains), and cross-region access.
  • Own golden images, templates, GPU drivers / CUDA, and the container / image registry, including versioned release and multi-region distribution. (Secondary scope.)
  • Build monitoring, capacity planning, and runbooks; eliminate single points of failure.
  • Partner with Compute (delivery), Network (storage fabric / RDMA), and Control Plane (provisioning / quota) teams.
How you will stand out:
  • 3+ years (Senior 6+) in storage engineering or platform infrastructure, with hands-on distributed / parallel file system experience.
  • Strong understanding of distributed file system internals — data / metadata separation, replication, consistency models, POSIX vs object semantics.
  • Proven experience integrating and operating distributed storage in production (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS, MinIO).
  • Performance tuning for high-throughput / parallel I/O; familiarity with NVMe, RDMA / RoCE storage networking, and caching is a strong plus.
  • Strong Linux systems depth and automation skills (Python / Go, CI / CD).
  • HPC / AI storage or multi-region storage experience a strong plus
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

#LI-ST1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Storage Engineer
Cloud Storage Engineer

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
Cloud Storage Integration Engineer
Cloud Storage Integration Engineer

United States Digital Space LLC • Singapore

On-site
SGD 150,000 - 190,000
Authentic culture and diverse team
Training and mentoring opportunities
Networking with industry pioneers
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Training and mentoring
Competitive benefits
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
Cloud Storage Integration Engineer (GPU Cloud)
Cloud Storage Integration Engineer (GPU Cloud)

United States Digital Space LLC • Singapore

On-site
SGD 150,000 - 190,000
Authentic culture and diverse team
Training and mentoring opportunities
Networking with industry pioneers
Cloud Senior DevOps Engineer
Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
SRE L1 Support/Cloud Platform Ops Engineers
SRE L1 Support/Cloud Platform Ops Engineers

Bitdeer Technologies Group • Singapore

On-site
SGD 4,000 - 7,000
Welfare benefits
Training & mentoring
Inclusive culture
SRE L1 Support/Cloud Platform Ops Engineers
SRE L1 Support/Cloud Platform Ops Engineers

Bitdeer • Singapore

On-site
SGD 52,000 - 78,000
AI Cloud Network Delivery Engineer
AI Cloud Network Delivery Engineer

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
Training and mentoring
Competitive welfare benefits
Global exposure