NVIDIA Storage Admin

PeopleStrong

Mumbai

On-site

INR 2,500,000 - 4,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PeopleStrong is seeking a BE/B-Tech with Computer Science or Electronics & Communications to lead a high-throughput storage infrastructure for large-scale AI training data, checkpoints, and inference artifacts. You will design/expand Lustre/BeeGFS and Object (S3) tiers aligned with AI dataflow and GPU clusters.

You will drive capacity planning, performance tuning, and automated snapshot/retention strategies, while leading P0/P1 incident response and conducting RCA for storage outages affecting

Qualifications

  • BE/B-Tech or equivalent in Computer Science or Electronics & Communications.
  • 7-12 years distributed storage experience.
  • Hands-on with Lustre, BeeGFS, Ceph, NVMe-oF and S3 in GPU environments.
  • Knowledge of multi-tenant encryption at rest and in transit, POSIX/ACLs and S3 policies.

Responsibilities

  • Design/expand Lustre/BeeGFS and Object (S3) tiers for AI dataflow and GPU workloads.
  • Capacity and performance planning; automate snapshots and checkpoint retention.
  • Lead P0/P1 incident response for enterprise-scale AI storage infrastructure.
  • Deep-dive troubleshooting for storage performance degradation and IO latency spikes.
  • Root cause analysis for storage-related outages affecting GPU training and inference.
  • Backup/DR planning and regular testing of restore, RPO/RTO.

Skills

Distributed storage
Incident management
Root cause analysis
Performance tuning
Capacity planning
Metadata handling

Education

BE/B-Tech or equivalent in Computer Science or Electronics & Communications

Tools

Lustre
BeeGFS
Ceph
NVMe-oF
S3
NFS
GPFS (IBM Spectrum Scale)

Job description

Bachelor of Technology (BTech)

Job Description

Job Purpose

Provide high?throughput, consistent storage tiers (Scratch/HPS + Object) for large?scale training data ingest, checkpoints, and inference artifacts.

  • Implementation

Design/expand Lustre/BeeGFS HPS; NVMe?oF and Object (S3) tiers; align with AI dataflow and GDS.

Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.

  • Operations

Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention.

Proactive detection of hot spots and metadata contention; schema for small?file handling.

  • Performance & Optimization

Tune RDMA paths, page cache, IO schedulers; validate end?to?end I/O profiles for LLM training/inference.

  • Reliability & Incident

Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads.

Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.

Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.

Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads.

Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split?brain handling.

  • Security & Compliance

Multi?tenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds.

Experience & Educational Requirement

BE/B-Tech or equivalent with Computer Science or Electronics & Communication

RELEVANT EXPERIENCE

  • 7-12 years distributed storage; hands?on with Lustre/BeeGFS/Ceph, NVMe?oF, and S3 in GPU environments.

Tools / Tech

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer, Storage and Data Protection
Engineer, Storage and Data Protection

AHEAD • Bengaluru

On-site
INR 1,800,000 - 2,400,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Staff Engineer
Staff Engineer

DDN • Pune District

On-site
INR 400,000 - 700,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Corporation • India

On-site
INR 3,500,000 - 7,000,000
NOC Technical Lead (L3)
NOC Technical Lead (L3)

Larsen & Toubro • Chennai District

On-site
INR 4,500,000 - 7,500,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 900,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000