Staff Engineer: Distributed Storage & AI Infra

Together Computer Inc

United States

Remote

USD 250,000 - 300,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Remote work flexibility

Job summary

Together AI is seeking a senior storage engineer to design and operate multi-petabyte storage systems for AI training and inference. You will integrate Vast, Weka, Ceph, and Lustre, build Kubernetes-native storage operators, and enable automated provisioning with strong multi-tenancy and performance isolation.

You will architect end-to-end data paths, optimize caching, and scale storage across thousands of nodes, delivering 10+ GB/s per GPU node while contributing to open-source and internal

Qualifications

  • 8+ years in storage engineering at multi-petabyte scale.
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters.
  • Deep Kubernetes and cloud-native storage experience in production environments.
  • Strong coding skills in Go and Python for production-grade systems.
  • BS/MS in Computer Science, Engineering, or equivalent practical experience.
  • History of technical leadership with improvements in performance, reliability, or cost.
  • Distributed storage systems expertise: Ceph, WekaFS, Lustre, Vast, GPFS, or similar.
  • Object storage experience with S3/MinIO/Ceph/R2 and performance optimization.
  • Kubernetes storage: CSI drivers, StatefulSets, PVs, storage operators, controllers.
  • Storage optimization for GPU workloads, RDMA/InfiniBand networking, TB/s throughput.
  • Programming: Go and Python for automation, operators, and tooling.
  • Infrastructure as Code: Terraform, Ansible, Helm, GitOps (ArgoCD).
  • Linux storage stack: Ext4, XFS, LVM, NVMe optimization, RAID.
  • Observability: Prometheus, Grafana, Thanos.

Responsibilities

  • Architect and implement the storage strategy and roadmap for Together AI as we scale GPU fleets.
  • Engineer and scale multi-petabyte AI/ML storage systems by integrating Vast, Weka, Ceph and optimizing costs via tiering.
  • Develop intelligent caching and tiered architectures to achieve extreme IOPS and throughput at GPU scale.
  • Tune storage isolation at L2/L3 network layers for secure multi-tenancy.
  • Code Kubernetes storage operators and controllers for automated provisioning and quota enforcement.
  • Engineer end-to-end data paths to achieve 10+ GB/s per GPU node and multi-tier caching for models and datasets.
  • Optimize data paths with benchmarking and profiling; contribute to open-source storage projects and internal tooling.

Skills

Go
Python
Kubernetes
Distributed storage
Cloud-native storage
Leadership
Linux storage stack

Education

BS/MS in Computer Science or Engineering

Tools

CSI drivers
Terraform
Ansible
Helm
GitOps (ArgoCD)

Job description

Together AI is seeking a senior storage engineer to design and operate multi-petabyte storage systems for AI training and inference. You will integrate Vast, Weka, Ceph, and Lustre, build Kubernetes-native storage operators, and enable automated provisioning with strong multi-tenancy and performance isolation.

You will architect end-to-end data paths, optimize caching, and scale storage across thousands of nodes, delivering 10+ GB/s per GPU node while contributing to open-source and internal

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Storage Architect: Petabyte-Scale AI Systems
Staff Storage Architect: Petabyte-Scale AI Systems

Cloudjobs • San Francisco (CA)

On-site
USD 200,000 - 260,000
Cash compensation
Equity compensation
Health coverage
+6
Staff Engineer, Distributed AI Storage Engine
Staff Engineer, Distributed AI Storage Engine

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with generous match
Catered lunch daily
+2
AI Storage Architect for High-Performance GPU Cloud
AI Storage Architect for High-Performance GPU Cloud

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 200,000
Senior Storage Engineer for AI Cloud Scale
Senior Storage Engineer for AI Cloud Scale

CoreWeave • Livingston (NJ)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Equity awards
+2
Staff Storage Engineer for AI Infrastructure
Staff Storage Engineer for AI Infrastructure

Socket.dev • Bellevue (WA)

On-site
USD 207,000 - 303,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+3
Senior AI Storage Engineer - Distributed Systems
Senior AI Storage Engineer - Distributed Systems

CoreWeave • Bellevue (WA)

On-site
USD 143,000 - 210,000
Medical, dental, vision insurance
Company 401(k) with match
Tuition Reimbursement
+3
Staff Storage Engineer: Scale Exabyte AI Storage
Staff Storage Engineer: Scale Exabyte AI Storage

EngineersOfAI • New York (NY)

On-site
USD 180,000 - 240,000
Senior Storage Engineer — Exabyte-Scale AI Storage
Senior Storage Engineer — Exabyte-Scale AI Storage

CoreWeave Europe • New York (NY)

On-site
USD 153,000 - 204,000
Medical insurance
Life Insurance
Disability insurance
+3
Storage Engineering Manager – AI Cloud Infra
Storage Engineering Manager – AI Cloud Infra

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Life Insurance
Disability insurance
+2
Storage Engineering Manager: Scale AI Storage 10x
Storage Engineering Manager: Scale AI Storage 10x

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
Equity awards
401(k) with employer match
+5