Storage Systems Engineer for Frontier AI

Prime Intellect

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Prime Intellect is building the open frontier AI stack and storage systems powering frontier workloads. You will design and operate storage architectures for training datasets, checkpoints, and model artifacts, balancing performance, durability, and cost as GPU clusters scale.

Join a hands-on, collaboration-heavy team working with Lustre, Ceph, and NVMe caching to support large GPU training and high-volume checkpoint workloads, with opportunities to contribute to open-source storage and scale

Qualifications

  • 3+ years building or operating production distributed storage systems.
  • Hands-on experience with at least one parallel or distributed filesystem or object storage platform (Lustre, BeeGFS, Ceph, GPFS).
  • Strong Linux administration and performance troubleshooting.
  • Experience automating infrastructure operations in Python, Go, Bash, or similar languages.
  • Understanding of storage failure modes, data integrity, replication, and recovery.

Responsibilities

  • Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows.
  • Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads.
  • Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads.
  • Build provisioning, capacity planning, lifecycle management, and operational automation for storage services.
  • Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets.
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices.
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks; collaborate with compute and networking teams.

Skills

Distributed storage
Linux administration
Scripting (Python/Go/Bash)
Storage systems knowledge
Performance troubleshooting

Tools

Lustre
BeeGFS
Ceph
GPFS

Job description

Prime Intellect is building the open frontier AI stack and storage systems powering frontier workloads. You will design and operate storage architectures for training datasets, checkpoints, and model artifacts, balancing performance, durability, and cost as GPU clusters scale.

Join a hands-on, collaboration-heavy team working with Lustre, Ceph, and NVMe caching to support large GPU training and high-volume checkpoint workloads, with opportunities to contribute to open-source storage and scale

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

Prime-Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Cluster Infra Architect for Frontier AI — Equity
Senior Cluster Infra Architect for Frontier AI — Equity

RadixArk • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Principal Software Engineer: Frontier AI Infra & Platform
Principal Software Engineer: Frontier AI Infra & Platform

NVIDIA • Santa Clara (CA)

On-site
USD 248,000 - 391,000
Equity
Benefits
Senior High-Performance Storage Architect - NVIS
Senior High-Performance Storage Architect - NVIS

NVIDIA Pty. Ltd • Santa Clara (CA)

On-site
AUD 250,000 - 347,000
Senior AI Storage Architect — High-Performance Linux & NVMe
Senior AI Storage Architect — High-Performance Linux & NVMe

NVIDIA Gruppe • California (MO)

On-site
USD 148,000 - 288,000
Equity
Benefits package
Senior AI Data Storage Engineer - Cloud-Native, Equity
Senior AI Data Storage Engineer - Cloud-Native, Equity

NVIDIA • North Carolina

On-site
USD 152,000 - 242,000
Equity compensation
Benefits
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime-Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
High-Performance Storage Architect, NVIS
High-Performance Storage Architect, NVIS

NVIDIA • California (MO)

On-site
USD 180,000 - 230,000