Sr. HPC ENGINEER

Cognizant

Hyderabad

On-site

INR 800,000 - 1,200,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cognizant is seeking seasoned HPC professionals to design, deploy, and maintain HPC clusters and storage systems in Hyderabad. You will optimize compute, storage, and networking performance while leading automation and monitoring efforts.

Role requires hands-on expertise with cluster managers, job schedulers, and enterprise storage; exposure to AI/ML workloads is a plus. You will work in a fast-paced environment delivering scalable HPC solutions.

Qualifications

  • Deep expertise in HPC platforms, including design, deployment and maintenance of HPC clusters.
  • Experience with storage systems such as Lustre, GPFS and enterprise storage solutions.
  • Proficiency in InfiniBand networking and high-speed interconnects.
  • Familiarity with cluster management tools like IBM LSF, Bright Cluster Manager, Altair Grid Manager.
  • Automation and scripting skills: Ansible, Chef, Cobbler, Bash, Python.
  • Experience with cloud HPC solutions such as AWS ParallelCluster.

Responsibilities

  • Operate and maintain HPC clusters on CentOS/RHEL and hardware like HPE and NVIDIA DGX.
  • Manage large-scale storage systems and optimize HPC storage performance.
  • Configure InfiniBand networking for low-latency, high-bandwidth communication.
  • Administer cluster management and job scheduling tools; optimize resource allocation.
  • Implement monitoring with Zabbix, Grafana, ELK; automate provisioning with Cobbler/Chef/Ansible and AWS ParallelCluster.
  • Perform performance benchmarking, tuning, and troubleshooting across compute, storage, and network layers.
  • Ensure security and compliance of the HPC environment.

Skills

Linux
Leadership
Communication
Problem-solving

Education

Bachelor's or Master's in CS/Engineering

Tools

Bright Cluster Manager
Altair Grid Manager
IBM LSF
Zabbix
Grafana
ELK Stack
Cobbler
Chef
Ansible
AWS ParallelCluster
Docker
Singularity
Lustre
GPFS
NVIDIA DGX
HPE

Job description

Role Overview

We are seeking seasoned professionals with deep expertise in HPC platforms, the ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.

Key Responsibilities
  • HPC Infrastructure Management
    • Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
    • Ensure optimal performance, scalability, and reliability of compute resources.
  • Storage Administration
    • Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
    • Implement data lifecycle management and optimize storage performance for HPC workloads.
  • Networking
    • Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
    • Troubleshoot network performance issues and ensure secure connectivity.
  • Cluster and Job Scheduling
    • Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
    • Optimize job scheduling and resource allocation for diverse workloads.
  • Monitoring and Automation
    • Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
    • Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
  • Performance Tuning & Troubleshooting
    • Conduct performance benchmarking and tuning for HPC workloads.
    • Diagnose and resolve hardware/software issues across compute, storage, and network layers.
  • Security & Compliance
    • Ensure HPC environment adheres to security best practices and compliance standards.
Required Skills & Qualifications
  • Technical Expertise
    • Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
    • Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
    • Proficiency in InfiniBand networking and high-speed interconnects.
    • Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
  • Automation & Scripting
    • Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
    • Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
  • Monitoring & Logging
    • Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
  • Soft Skills
    • Strong problem-solving and analytical skills.
    • Ability to work in a fast-paced environment and lead technical teams.
    • Excellent communication and documentation skills.
Preferred Qualifications
  • Exposure to AI/ML workloads on HPC clusters.
  • Experience with containerization (Docker, Singularity) in HPC environments.
  • Knowledge of security hardening for HPC systems.
Education
  • Bachelor's or Master's degree in Computer Science, Engineering, or related field.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Engineer
Senior HPC Engineer

Netweb Technologies India Ltd. • Faridabad District

On-site
INR 3,500,000 - 5,500,000
HPC Senior System Integrator/System Administrator
HPC Senior System Integrator/System Administrator

GBB • Mumbai

On-site
INR 1,000,000 - 1,500,000
Senior HPC Engineer
Senior HPC Engineer

Binaire Private Limited • New Delhi

On-site
INR 1,500,000 - 2,500,000
Opportunity to influence hardware selection
Ownership of high-performance compute infrastructure
Linux System Administrator
Linux System Administrator

SISL Global • Chennai District

On-site
INR 800,000 - 1,200,000
HPC Admin
HPC Admin

Randstad Digital • Bengaluru

On-site
INR 900,000 - 1,500,000
Lead HPC Engineer
Lead HPC Engineer

Clovertex • Hyderabad

On-site
INR 2,000,000 - 3,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA • Bengaluru

On-site
INR 2,500,000 - 4,000,000
HPC Engineer
HPC Engineer

Whiteblue • Chennai

On-site
INR 1,500,000 - 2,500,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 900,000
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]

SanDisk • Bengaluru

On-site
INR 3,000,000 - 5,000,000