Cloud Storage Integration Engineer (Storage / Image / Registry)

Bitdeer (NASDAQ: BTDR)

Austin (TX)

On-site

USD 140,000 - 210,000

Full time

29 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer Technologies Group is seeking an experienced Storage/Platform Engineer to design and integrate distributed file systems for the GPU cloud, focusing on high-throughput AI training and inference workloads.

You will own end-to-end storage integration, tune performance, and architect multi-region data delivery. Advanced Linux skills and scripting are essential, with HPC/AI storage experience a strong plus.

Qualifications

  • 3+ years (Senior 6+) in storage engineering or platform infrastructure.
  • Strong understanding of distributed file system internals — data/metadata separation, replication, consistency models.
  • Proven experience integrating and operating distributed storage in production (Ceph, Lustre, GPFS/Spectrum Scale, BeeGFS, JuiceFS, MinIO).
  • Performance tuning for high-throughput/parallel I/O; familiarity with NVMe, RDMA/RoCE storage networking, and caching is a strong plus.
  • Strong Linux systems depth and automation skills (Python/Go, CI/CD).
  • HPC/AI storage or multi-region storage experience a strong plus.

Responsibilities

  • Design and integrate distributed/parallel file systems (Ceph, Lustre, GPFS/Spectrum Scale, BeeGFS, JuiceFS) into the GPU cloud, optimized for AI training/inference I/O patterns.
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform.
  • Tune storage throughput and latency for large-scale parallel access; benchmark across GPU SKUs and workloads.
  • Architect multi-region storage: data locality, replication/consistency, durability, and cross-region access.
  • Own golden images, templates, GPU drivers/CUDA, and the container/image registry, including versioned release and multi-region distribution.
  • Build monitoring, capacity planning, and runbooks; eliminate single points of failure.

Skills

Distributed storage
Linux systems
Python/Go scripting
CI/CD automation

Tools

Ceph
Lustre
GPFS/Spectrum Scale
BeeGFS
JuiceFS
MinIO
NVMe
RDMA

Job description

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

GPU training and inference at 10,000+ GPU, multi-region scale depend on high-throughput, low-latency storage that can sustain massive parallel I/O. We are looking for an engineer who deeply understands distributed file systems and can integrate distributed / parallel storage systems into our GPU cloud — covering performance, multi-tenancy, and reliability — while also owning the image / driver / registry pipeline on the node-delivery critical path.

Key Responsibilities
  • Design and integrate distributed / parallel file systems (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS) into the GPU cloud, optimized for AI training / inference I/O patterns.
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform / control plane.
  • Tune storage throughput and latency for large-scale parallel access (dataset loading, checkpointing); benchmark across GPU SKUs and workloads.
  • Architect multi-region storage: data locality, replication / consistency, durability (failure domains), and cross-region access.
  • Own golden images, templates, GPU drivers / CUDA, and the container / image registry, including versioned release and multi-region distribution. (Secondary scope.)
  • Build monitoring, capacity planning, and runbooks; eliminate single points of failure.
  • Partner with Compute (delivery), Network (storage fabric / RDMA), and Control Plane (provisioning / quota) teams.
Job Requirement
  • 3+ years (Senior 6+) in storage engineering or platform infrastructure, with hands-on distributed / parallel file system experience.
  • Strong understanding of distributed file system internals — data / metadata separation, replication, consistency models, POSIX vs object semantics.
  • Proven experience integrating and operating distributed storage in production (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS, MinIO).
  • Performance tuning for high-throughput / parallel I/O; familiarity with NVMe, RDMA / RoCE storage networking, and caching is a strong plus.
  • Strong Linux systems depth and automation skills (Python / Go, CI / CD).
  • HPC / AI storage or multi-region storage experience a strong plus.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Storage Integration Engineer (Storage / Image / Registry)
Cloud Storage Integration Engineer (Storage / Image / Registry)

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 170,000 - 210,000
Sr. GPU Cloud Storage Solutions Expert (SRE SME)
Sr. GPU Cloud Storage Solutions Expert (SRE SME)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 250,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
AI Storage Solutions Expert
AI Storage Solutions Expert

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 200,000
AI Storage Solutions Expert
AI Storage Solutions Expert

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
GPU Cloud Storage Engineer — Distributed FS & Registry
GPU Cloud Storage Engineer — Distributed FS & Registry

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
GPU Cloud Storage Engineer — Multi-Region AI I/O
GPU Cloud Storage Engineer — Multi-Region AI I/O

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 210,000