Staff Storage Platform Engineer (AI Storage) - Radian Arc

Jobgether

France

Sur place

EUR 120 000 - 180 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Remote-friendly Europe
Global team
Career growth
Flexible hours

Résumé du poste

Radian Arc in France seeks a Staff Storage Platform Engineer (AI Storage) to architect and own high-performance storage for large-scale AI workloads. You will design, build, and operate storage systems spanning edge and core GPU deployments, with strong emphasis on throughput, latency, resilience, and cost efficiency.

You will lead platform integration with Kubernetes CSI, S3-compatible object storage, and multi-cluster deployments, mentor engineers, and define reusable patterns while shaping

Qualifications

  • Strong hands-on experience designing and operating distributed storage systems in HPC/AI environments.
  • Proven experience designing storage for large-scale AI training, fine-tuning, or inference.
  • Deep understanding of how AI workloads affect storage performance, latency, and data locality.
  • Experience with StorPool, VAST Data, Weka, local NVMe, distributed filesystems, and S3/object storage.
  • Linux I/O stack, NVMe devices, storage fabrics, and high-performance data paths.
  • Kubernetes CSI storage integrations and multi-tenant storage architectures.
  • Programming/scripting for automation (Python or Bash).
  • Observability and telemetry to diagnose performance and reliability issues.
  • Ability to lead complex infra initiatives and mentor engineers.

Responsabilités

  • Design scalable storage architectures for edge and core GPU deployments with high throughput and low latency.
  • Optimize storage for distributed training, fine-tuning, and inference workloads.
  • Design storage for inference platforms and KV-cache persistence to scale across large GPU clusters.
  • Engineer storage-to-GPU data paths using GPU Direct Storage, RDMA, NVMe-oF, SPDK.
  • Integrate block, object, and shared file storage into Kubernetes with CSI.
  • Contribute to S3-compatible object storage and multi-cluster storage platforms.
  • Lead benchmarking, capacity planning, incident response, and root-cause analysis.
  • Own storage initiatives end-to-end from architecture to production rollout.
  • Improve observability, automation, runbooks, and day-2 operations.
  • Mentor engineers and influence platform architecture and roadmap.

Connaissances

Distributed storage
AI infrastructure
AI storage knowledge
Storage technologies
Linux and systems
Kubernetes
AI data paths
Distributed inference
Troubleshooting
Automation
Observability
Technical leadership
Systems thinking
Communication
Ownership
Mentoring

Description du poste

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Storage Platform Engineer (AI Storage) - Radian Arc based in France.

This is aStaff-level opportunityto shape thestorage architecture poweringlarge-scale GPUand AI infrastructureacross edge andcore environments.
You will design, build, and operatehigh-performance storagesystems supporting training, fine-tuning, anddistributed inference workloads.
The role covershyperconverged, local NVMe, and disaggregated storagearchitectures, with astrong focus onthroughput, latency, resilience, and cost efficiency.
You will workat the intersectionof storage, GPUs, networking, Kubernetes, andAI platform engineering, ensuring storage neverbecomes a bottleneckfor compute.
As the primarystorage specialist, youwill combine architecturalownership with hands-on engineering, troubleshooting, performance optimization, anddeployment.
You will alsodefine reusable standards, influence long-termplatform direction, andmentor engineers acrossadjacent infrastructuredomains.
This is anideal environment fora senior storageexpert who wantssignificant technical ownershipand direct impacton next-generationAI infrastructure.

Accountabilities
  • Storage architecture:Design scalable storage architectures for edge and core GPU deployments, covering hyperconverged platforms such as StorPool, local NVMe, and disaggregated systems such as VAST Data and Weka. Define reference architectures, reusable design patterns, fault domains, lifecycle strategies, and scaling approaches while balancing throughput, latency, resilience, data locality, operability, and cost.
  • AI workload optimization:Optimize storage for distributed training, fine-tuning, and inference workloads, including large dataset ingestion, model artifact distribution, checkpointing, and high-concurrency access. Establish realistic performance baselines and ensure storage architecture aligns with actual GPU workload behavior.
  • Distributed inference:Design storage architectures supporting inference platforms such as NVIDIA Dynamo, llm-d, or similar systems. Optimize model distribution, token-generation data paths, and KV-cache persistence and retrieval so infrastructure can scale efficiently across large GPU clusters without storage becoming a throughput or latency bottleneck.
  • High-performance data paths:Engineer efficient storage-to-GPU data paths using technologies such as GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK. Investigate and tune performance across hardware, networking, operating systems, filesystems, storage layers, and distributed workloads.
  • Platform integration:Integrate block, object, and shared file storage into Kubernetes and platform orchestration systems. Implement and maintain CSI integrations, support multi-tenant storage architectures, and define standards for storage integration across different deployment models.
  • Distributed storage:Contribute to large-scale storage platforms, including S3-compatible object storage, distributed file systems, and block storage. Design systems with clear operational boundaries, resilience models, scaling paths, and reusable operating patterns across multi-cluster and multi-site environments.
  • Performance and reliability:Lead storage benchmarking, capacity planning, performance investigations, incident response, and root-cause analysis. Establish measurable standards for throughput, latency consistency, recovery behavior, reliability, and operational maturity.
  • Engineering delivery:Own storage initiatives end to end, from architecture and validation through production rollout. Validate BOMs, topology decisions, node profiles, and deployment assumptions while ensuring changes are introduced safely with minimal customer impact.
  • Operational excellence:Improve storage observability, automation, runbooks, lifecycle management, and day-2 operations. Turn recurring incidents and operational pain points into durable engineering improvements and standardized practices.
  • Technical leadership:Act as the primary storage design authority, influencing platform architecture and roadmap decisions across compute, networking, DevOps, infrastructure, and operations. Communicate technical trade-offs clearly, mentor adjacent engineers, and raise the organization’s expertise in AI storage.
Requirements
  • Distributed storage expertise:Strong hands-on experience designing and operating distributed storage systems in high-performance computing, AI, GPU, or similarly demanding environments.
  • AI infrastructure experience:Proven experience designing storage architectures for large-scale AI training, fine-tuning, or inference, including dataset distribution, model artifacts, checkpointing, and high-concurrency data access.
  • AI storage knowledge:Deep understanding of how AI workload characteristics affect storage throughput, latency, concurrency, data locality, checkpoint recovery, and serving performance.
  • Storage technologies:Hands-on experience with technologies such as Weka, VAST Data, StorPool, local NVMe, distributed filesystems, S3-compatible object storage, block storage, and/or comparable enterprise storage platforms.
  • Linux and systems expertise:Strong knowledge of the Linux storage and I/O stack, storage hardware, NVMe devices, storage fabrics, and high-performance data paths.
  • Kubernetes:Familiarity with Kubernetes storage integrations, particularly CSI, and experience integrating storage into containerized or orchestrated platforms.
  • AI data paths:Practical knowledge of GPU Direct Storage, RDMA/RoCE, NVMe-oF, SPDK, and techniques for minimizing unnecessary data movement between storage and GPU compute.
  • Distributed inference:Experience with storage requirements for inference orchestration and model-serving environments, including model distribution and KV-cache persistence or retrieval, is highly valuable.
  • Troubleshooting:Ability to diagnose complex cross-layer issues involving storage hardware, networking, Linux kernels and I/O paths, filesystems, object/block storage, Kubernetes, and distributed workloads.
  • Automation:Strong Python and/or Bash skills, with experience applying software engineering practices to infrastructure automation, validation, lifecycle management, and operational tooling.
  • Observability:Experience designing or operating storage observability systems and using metrics and telemetry to identify performance, reliability, and capacity issues.
  • Technical leadership:Demonstrated ability to lead complex infrastructure initiatives, establish architectural standards, and influence multiple teams without relying on formal management authority.
  • Systems thinking:Ability to balance performance, scalability, reliability, operability, deployment complexity, and cost when making architecture decisions.
  • Communication and collaboration:Comfortable working with compute, networking, platform, DevOps, operations, deployment teams, vendors, and other technical stakeholders.
  • Ownership:Able to combine Staff-level strategic thinking with hands-on execution, particularly in a lean or fast-scaling environment where processes and standards are still being established.
  • Mentoring:Strong ability to share knowledge, guide engineers in adjacent domains, and raise the technical bar across the broader infrastructure organization.
Benefits
  • Attractive compensation package aligned with your expertise and experience.
  • Opportunity to play a foundational role in shaping a next-generation AI storage platform.
  • Significant architectural ownership and direct influence over long-term infrastructure strategy.
  • Exposure to cutting-edge GPU, AI inference, distributed storage, and high-performance data technologies.
  • International and diverse working environment with strong flexibility.
  • Remote-friendly work model across Europe.
  • Opportunity to join a fast-growing scale-up with an ambitious technology mission.
  • Broad cross-functional exposure across infrastructure, compute, networking, platform engineering, and operations.
  • Strong career growth potential as the infrastructure organization expands.
  • Inclusive environment committed to equal opportunity and professional development.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior Sales Engineer - Strategic AI
Senior Sales Engineer - Strategic AI

DDN • Paris

Sur place
EUR 70 000 - 110 000
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)

Jobgether • France

Hybride
EUR 120 000 - 180 000
Flexible hybrid work
Career growth opportunities
Autonomy over networking architecture
+2
Senior AI Storage Platform Architect
Senior AI Storage Platform Architect

Jobgether • France

Sur place
EUR 120 000 - 180 000
Remote-friendly Europe
Global team
Career growth
+1
Senior Sales Engineer
Senior Sales Engineer

DDN • Paris

Sur place
EUR 85 000 - 120 000
Data Center Engineer
Data Center Engineer

Thor • Paris

Sur place
EUR 70 000 - 100 000
Data Center Engineer
Data Center Engineer

Trust In SODA • Paris

Sur place
EUR 60 000 - 90 000
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA Gruppe • Courbevoie

Sur place
EUR 120 000 - 180 000
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA • Courbevoie

Sur place
EUR 90 000 - 140 000
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA Corporation • France

À distance
EUR 90 000 - 120 000
Senior AI Compute Engineer
Senior AI Compute Engineer

NVIDIA • France

Sur place
EUR 90 000 - 150 000