Cloud Storage Integration Engineer (Storage / Image / Registry)

Bitdeer (NASDAQ: BTDR)

San Jose (CA)

On-site

USD 150,000 - 190,000

Full time

48 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Bitdeer Technologies Group seeks an experienced storage infra engineer to design and integrate distributed/parallel file systems into its GPU cloud for AI training and inference at scale.

You will own end-to-end storage integration, tune throughput/latency, architect multi-region data distribution, and manage images, drivers, and registries while collaborating with Compute, Network, and Control Plane teams.

Qualifications

  • 3+ years in storage engineering or platform infrastructure, senior level 6+
  • Deep understanding of distributed file system internals—data/metadata, replication, consistency
  • Production experience with distributed storage (Ceph, Lustre, GPFS/Spectrum Scale, BeeGFS, JuiceFS, MinIO)
  • Performance tuning for high-throughput parallel I/O; NVMe / RDMA knowledge a plus
  • Strong Linux depth and automation skills (Python / Go, CI/CD)
  • HPC/AI storage or multi-region storage experience is a strong plus

Responsibilities

  • Design and integrate distributed/parallel file systems into GPU cloud for AI training/inference
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quotas
  • Tune storage throughput/latency for large-scale parallel access across GPU SKUs
  • Architect multi-region storage with data locality, replication, durability, cross-region access
  • Own golden images, templates, GPU drivers/CUDA, and container/image registry distribution
  • Build monitoring, capacity planning, and runbooks; partner with Compute/Network/Control Plane teams

Skills

Distributed file systems
Python
Go
Linux
CI/CD
Performance tuning

Tools

Ceph
Lustre
GPFS/Spectrum Scale
BeeGFS
JuiceFS
MinIO

Job description

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

GPU training and inference at 10,000+ GPU, multi-region scale depend on high-throughput, low-latency storage that can sustain massive parallel I/O. We are looking for an engineer who deeply understands distributed file systems and can integrate distributed / parallel storage systems into our GPU cloud — covering performance, multi-tenancy, and reliability — while also owning the image / driver / registry pipeline on the node-delivery critical path.

Key Responsibilities
  • Design and integrate distributed / parallel file systems (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS) into the GPU cloud, optimized for AI training / inference I/O patterns.
  • Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform / control plane.
  • Tune storage throughput and latency for large-scale parallel access (dataset loading, checkpointing); benchmark across GPU SKUs and workloads.
  • Architect multi-region storage: data locality, replication / consistency, durability (failure domains), and cross-region access.
  • Own golden images, templates, GPU drivers / CUDA, and the container / image registry, including versioned release and multi-region distribution. (Secondary scope.)
  • Build monitoring, capacity planning, and runbooks; eliminate single points of failure.
  • Partner with Compute (delivery), Network (storage fabric / RDMA), and Control Plane (provisioning / quota) teams.
Job Requirement
  • 3+ years (Senior 6+) in storage engineering or platform infrastructure, with hands-on distributed / parallel file system experience.
  • Strong understanding of distributed file system internals — data / metadata separation, replication, consistency models, POSIX vs object semantics.
  • Proven experience integrating and operating distributed storage in production (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS, MinIO).
  • Performance tuning for high-throughput / parallel I/O; familiarity with NVMe, RDMA / RoCE storage networking, and caching is a strong plus.
  • Strong Linux systems depth and automation skills (Python / Go, CI / CD).
  • HPC / AI storage or multi-region storage experience a strong plus.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Storage Integration Engineer (Storage / Image / Registry)
Cloud Storage Integration Engineer (Storage / Image / Registry)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 210,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 170,000 - 210,000
Sr. GPU Cloud Storage Solutions Expert (SRE SME)
Sr. GPU Cloud Storage Solutions Expert (SRE SME)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 250,000
Senior AI Storage Infrastructure Engineer
Senior AI Storage Infrastructure Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
AI Storage Solutions Expert
AI Storage Solutions Expert

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 200,000
AI Storage Solutions Expert
AI Storage Solutions Expert

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
GPU Cloud Storage Engineer — Distributed FS & Registry
GPU Cloud Storage Engineer — Distributed FS & Registry

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 150,000 - 190,000
GPU Cloud Storage Engineer — Multi-Region AI I/O
GPU Cloud Storage Engineer — Multi-Region AI I/O

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 210,000