Engineer Storage And Data Protection

AHEAD

Gurugram District

On-site

INR 1,500,000 - 2,100,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AHEAD is seeking an experienced HPC storage engineer to provide enterprise-level operational support for Managed Services customers in India. You will administer distributed filesystems, optimize performance for AI training, and build automation for provisioning, monitoring, and lifecycle management.

The role requires strong Linux administration, experience with Lustre/GPFS/Ceph, and knowledge of Slurm and Kubernetes.

Qualifications

  • 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering.
  • Bachelors degree or equivalent Information Systems or related field.
  • Strong experience with Linux systems administration.
  • Hands-on experience configuring, managing, and tuning distributed or parallel filesystems.
  • Experience tuning storage for performance-sensitive workloads.
  • Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes
  • Familiarity with high-speed interconnects such as InfiniBand or RDMA
  • Ability to troubleshoot complex issues across storage, compute, and networking layers
  • Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments
  • Experience with machine learning or data science workflows in HPC environments
  • Managed Services or consulting experience
  • Strong background with customer service
  • High level problem-solving and communication skills
  • Managed Services or consulting experience

Responsibilities

  • Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities.
  • Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast.
  • Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference.
  • Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations.
  • Plan and perform maintenance activities.
  • Assess customer environments for performance and design issues and propose resolutions.
  • Work across technical teams to troubleshoot complex infrastructure issues.
  • Create and maintain detailed documentation.
  • Serve as a subject matter expert and escalation point for storage technologies.
  • Work with vendors to resolve storage issues.
  • Communicate with customers and internal team with transparency.
  • Support data movement workflows including ingest, replication, caching, tiering, and archiving.
  • Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics.
  • Partner with infrastructure, platform, and research teams to support production AI/HPC workloads.
  • Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency.
  • Communicate with customers and internal team with transparency.
  • Participate in on-call rotation

Skills

Linux administration
Distributed filesystems
HPC infrastructure
Kubernetes
Slurm
RDMA
Performance tuning
Customer service
Communication skills
Problem solving

Education

Bachelors degree in Information Systems

Tools

Terraform
Ansible
Helm
GitOps
Prometheus
Grafana
Ceph
MinIO

Job description

Job Description:

Key Responsibilities
  • Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities
  • Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast
  • Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference
  • Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations
  • Plan and perform maintenance activities
  • Assess customer environments for performance and design issues and propose resolutions
  • Work across technical teams to troubleshoot complex infrastructure issues
  • Create and maintain detailed documentation
  • Serve as a subject matter expert and escalation point for storage technologies
  • Work with vendors to resolve storage issues
  • Communicate with customers and internal team with transparency
  • Support data movement workflows including ingest, replication, caching, tiering, and archiving
  • Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics
  • Partner with infrastructure, platform, and research teams to support production AI/HPC workloads
  • Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency
  • Communicate with customers and internal team with transparency
  • Participate in on-call rotation
Required Qualifications
  • 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering
  • Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education
  • Strong experience with Linux systems administration
  • Hands-on experience configuring, managing, and tuning distributed or parallel filesystems
  • Experience tuning storage for performance-sensitive workloads
  • Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes
  • Familiarity with high-speed interconnects such as InfiniBand or RDMA
  • Ability to troubleshoot complex issues across storage, compute, and networking layers
  • Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments
  • Experience with machine learning or data science workflows in HPC environments
  • Managed Services or consulting experience
  • Strong background with customer service
  • High level problem-solving and communication skills
  • Strong oral and written communications skills
  • Managed Services or consulting experience
Preferred Qualifications
  • Experience supporting storage solutions for GPU clusters and AI/ML workflows
  • Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway
  • Experience with Terraform, Ansible, Helm, or GitOps workflows
  • Knowledge of observability platforms such as Prometheus and Grafana
  • Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems
  • Experience with machine learning or data science workflows in HPC environments
  • Scripting or programming experience with Python and Bash
  • Related Storage certifications are a bonus

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Requirements:

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer, Storage and Data Protection
Engineer, Storage and Data Protection

AHEAD • Gurugram District

On-site
INR 2,800,000 - 4,500,000
Engineer, Storage and Data Protection
Engineer, Storage and Data Protection

AHEAD • Bengaluru

On-site
INR 1,500,000 - 2,700,000
Storage Systems Engineer (C/C++, File System, Storage)
Storage Systems Engineer (C/C++, File System, Storage)

Hewlett Packard Enterprise • Bengaluru

On-site
INR 2,800,000 - 5,400,000
Career development programs
Inclusive culture
Flexible work policies
Sr. HPC ENGINEER
Sr. HPC ENGINEER

Cognizant • Hyderabad

On-site
INR 800,000 - 1,200,000
HPC Storage Engineer
HPC Storage Engineer

SourcingXPress • Hyderabad

On-site
INR 5,000,000 - 9,000,000
Storage Systems Engineer
Storage Systems Engineer

Hewlett Packard Enterprise • Bengaluru

Hybrid
INR 1,200,000 - 1,600,000
Comprehensive health benefits
Career development programs
Flexible work schedules
Storage Systems Engineer (C/C++, File System, Storage)
Storage Systems Engineer (C/C++, File System, Storage)

Hewlett Packard Enterprise Development LP • India

On-site
INR 2,500,000 - 4,200,000
Health & wellbeing
Personal & professional development
Unconditional inclusion
Storage Systems Engineer (C/C++, File System, Storage)
Storage Systems Engineer (C/C++, File System, Storage)

Hewlett Packard Enterprise Development LP • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Health and wellbeing benefits
Career development programs
Inclusive culture
Senior DevOps Engineer – Storage Platforms
Senior DevOps Engineer – Storage Platforms

Jobtailor • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Senior Sales Engineer
Senior Sales Engineer

DDN • Delhi

On-site
INR 1,200,000 - 2,400,000