Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]

SanDisk

Bengaluru

On-site

INR 3,000,000 - 5,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SanDisk is seeking an experienced HPC Administrator to design and deploy high-performance computing clusters in Bengaluru. You will configure hardware, networking, and storage for performance and reliability, perform routine maintenance, monitor capacity, and provide user support.

The role requires strong Linux administration, experience with Slurm/LSF/Grid Engine, and knowledge of storage systems (NetApp/Isilon).

Qualifications

  • Bachelor’s degree in Computer Science, IT or related field.
  • 8+ years of experience as an HPC Administrator or similar role.
  • Experience with Slurm, NC, LSF or Grid Engine in production HPC.
  • Strong Linux/Unix admin and shell scripting skills.
  • Experience with NFS/storage and backup management in HPC.
  • Familiarity with networking concepts (TCP/IP, VLANs, InfiniBand).
  • Excellent analytical, problem-solving, and cross-team collaboration abilities.

Responsibilities

  • Design and deploy high-performance computing clusters and systems.
  • Configure hardware, networks and storage to optimize performance and reliability.
  • Perform routine maintenance, updates, patches, and upgrades.
  • Monitor system performance, capacity planning, and bottlenecks.
  • Provide technical support and troubleshooting to HPC users.
  • Deliver training sessions on best practices and usage guidelines.
  • Implement security controls and data protection for HPC infrastructure.
  • Ensure regulatory compliance and internal policies for HPC operations.
  • Create and maintain documentation and troubleshooting guides.
  • Generate reports on performance and usage for management.

Skills

HPC administration
Linux administration
Networking
Shell scripting
Problem solving
Cross-functional collaboration

Education

Bachelor’s degree in CS/IT or related field

Tools

Slurm
LSF/Grid Engine
NFS/Storage (NetApp/Isilon)
Docker
Kubernetes
Ansible
UCS servers

Job description

"
Job Description


  1. Designing and deploying high-performance computing clusters and systems based on organizational requirements and industry best practices.

  2. Configuring hardware components, network and storage systems to optimize performance and reliability.

  3. Performing routine maintenance tasks such as software updates, patches, and system upgrades to ensure optimal performance and security.

  4. Monitoring system performance, resource utilization, and capacity planning to proactively address potential issues and bottlenecks.

  5. Providing technical support and troubleshooting assistance to users of the HPC systems.

  6. Developing and delivering training sessions to educate users on best practices, usage guidelines, and efficient utilization of HPC resources.

  7. Implementing and maintaining security protocols, access controls, and data protection measures to safeguard HPC infrastructure and sensitive data.

  8. Ensuring compliance with relevant regulatory requirements and organizational policies related to HPC operations.

  9. Creating and maintaining comprehensive documentation including system configurations, operational procedures, and troubleshooting guides.

  10. Generating regular reports on system performance, usage statistics, and operational metrics for management and stakeholders.


Qualifications


  • Bachelor’s degree in computer science, Information Technology, or a related field (or equivalent work experience).

  • Proven experience (8+ years) as an HPC Administrator or in a similar role managing HPC systems in a production environment.

  • Proficiency in configuring and managing HPC cluster software such as Slurm, NC, LSF or Grid Engine.

  • Strong knowledge of Linux/Unix system administration and shell scripting.

  • Experience with NFS and storage (NetApp/ISILON) and backup management in HPC environments.

  • Familiarity with networking principles, including TCP/IP, VLANs, and InfiniBand.

  • Excellent analytical and problem-solving skills with the ability to troubleshoot complex issues independently.

  • Strong communication skills and the ability to collaborate effectively with cross-functional and cross geography teams and end-users.


Preferred Skills


  • Bachelor’s degree in computer science, Engineering, or a related discipline.

  • Experience in HPC technologies (e.g., HPC Systems Professional, Cray Certified System Administrator).

  • Knowledge with containerization technologies (e.g., Docker, Singularity) and workload orchestration frameworks (e.g., Kubernetes) is a plus.

  • Knowledge of scripting languages like shell/Ansible commonly used in unix admin will be a plus.

  • Knowledge of Dell/CISCO UCS servers in HPC environments.

  • Semiconductor domain experience is a must.


Additional Information

Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at jobs.accommodations@sandisk.com to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

"
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Engineer - High Performance Computing (HPC) and Unix Systems
Principal Engineer - High Performance Computing (HPC) and Unix Systems

SanDisk India Device Design Centre Pvt.Ltd • Bengaluru

On-site
INR 2,000,000 - 2,800,000
Principal Engineer, Data Engineering (HPC cluster software such as Slurm, NC, LSF or Grid Engin[...]
Principal Engineer, Data Engineering (HPC cluster software such as Slurm, NC, LSF or Grid Engin[...]

Sandisk • Bengaluru

On-site
INR 3,500,000 - 6,000,000
HPC Engineer
HPC Engineer

Sandisk • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Senior HPC Engineer
Senior HPC Engineer

Netweb Technologies India Ltd. • Faridabad District

On-site
INR 1,500,000 - 2,100,000
Linux System Administrator
Linux System Administrator

SISL Global • Chennai District

On-site
INR 800,000 - 1,200,000
HPC Senior System Integrator/System Administrator
HPC Senior System Integrator/System Administrator

GBB • Mumbai

On-site
INR 1,000,000 - 1,500,000
Storage Systems Engineer (C/C++, File System, Storage)
Storage Systems Engineer (C/C++, File System, Storage)

Hewlett Packard Enterprise Development LP • India

On-site
INR 2,500,000 - 4,200,000
Health & wellbeing
Personal & professional development
Unconditional inclusion
Storage Systems Engineer (C/C++, File System, Storage)
Storage Systems Engineer (C/C++, File System, Storage)

Hewlett Packard Enterprise Development LP • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Health and wellbeing benefits
Career development programs
Inclusive culture
Engineer II Systems (HPC Engineer)
Engineer II Systems (HPC Engineer)

Microchip Technology Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
HPC Engineer
HPC Engineer

Yotta Data Services Private Limited • Mumbai

On-site
INR 400,000 - 700,000