HPC Engineer

Esconet Technologies

New Delhi

On-site

INR 1,200,000 - 1,700,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Esconet Technologies is seeking an experienced HPC System Integrator/Administrator to lead the integration, delivery, and management of open-source HPC products. The role covers cluster administration, validation, and multi-site Linux HPC server operations.

You will work with Intel/AMD CPUs, NVIDIA GPUs, storage, InfiniBand, and Linux software, focusing on high-performance computing stack design and AI framework integration.

Qualifications

  • Bachelor's degree with 3+ years of relevant work experience or equivalent combination
  • 3+ years of experience with software development in Linux
  • 3+ years of experience with HPC clusters and systems integration
  • Ability to manage the AI stack up to the framework level
  • Experience with GPU cluster capabilities and storage integration (Lustre/BeeGFS)
  • Working knowledge of object-based storage, networking, InfiniBand, solution design and documentation

Responsibilities

  • Install, configure, fine-tune, and troubleshoot multi-vendor, multi-site Linux HPC servers
  • Build and deploy open-source software as well as vendor software
  • Diagnose and resolve system operational issues quickly and effectively
  • Verify full operation of systems, including network, systems and storage performance
  • Configure scheduling and queuing systems
  • Assist technical support teams with questions from customers
  • Coordinate with vendors to resolve hardware and software problems
  • Document system administration procedures (wikis)
  • Maintain and monitor the security of HPC systems and servers
  • Design and configure HPC/Kubernetes clusters with NVIDIA GPU support including MIG and AI Factory deployments
  • Deploy and manage Lustre storage cluster (OSS and MDS)
  • Build and maintain relationships with coworkers, managers and clients
  • Travel on-site for installation/maintenance as needed

Skills

Parallel processing
MPI/OpenMP
AI stack
GPU cluster
Networking
Scripting
Ansible
Flask dashboards

Education

Bachelor's degree in CS/CE/Computational Science or related field

Tools

Lustre
BeeGFS
InfiniBand
GPFS
Kubernetes

Job description

Experience: 5+ Years

Esconet Technologies is looking for a highly motivated HPC System Integrator/Administrator with a strong passion for cluster administration, system integration, and validation of HPC clusters. In this role, you will influence the overall integration, delivery, and management of largely open-source HPC products and solutions — spanning Intel and AMD processors, NVIDIA GPUs, storage, InfiniBand, and Linux software.

A solid understanding of parallel processing (problem decomposition and work distribution), parallel programming (MPI, OpenMP), and computer architecture is essential for this role.

Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, Computational Science, equivalent mathematical sciences, or a related field, with 3+ years of relevant work experience; or an equivalent combination of education, training, and experience
  • 3+ years of experience with software development in Linux
  • 3+ years of experience with HPC clusters and systems integration
  • Ability to manage the AI stack up to the framework level
  • Experience with GPU cluster capabilities and storage integration (LUSTRE/BeeGFS, cluster storage)
  • Working knowledge of object-based storage, networking, InfiniBand, solution designing, and technical documentation
Key Responsibilities
  • Install, configure, fine-tune, and troubleshoot multi-vendor, multi-site Linux HPC servers
  • Build and deploy open-source software as well as vendor/partner software
  • Diagnose and resolve system operational issues quickly and effectively
  • Verify full operation of systems, including network, systems, and storage performance
  • Configure scheduling and queuing systems
  • Assist technical support teams with questions and issues encountered by customers
  • Coordinate with vendors to resolve hardware and software problems
  • Document system administration procedures for routine and complex tasks (wikis)
  • Maintain and monitor the security of HPC systems and servers
  • Design and configure HPC/Kubernetes clusters with NVIDIA GPU support, including MIG (Multi-Instance GPU) and AI Factory deployments
  • Deploy and manage cluster file systems such as Lustre, including OSS (Object Storage Server) and MDS (Metadata Server) node configuration
  • Build and maintain effective working relationships with coworkers, managers, and clients
  • Travel as needed for on-site cluster installation or maintenance (limited)
Desired Skills & Experience
  • Building, configuring, and administering Linux distributions — Rocky Linux, Ubuntu, CentOS, RHEL, and SUSE
  • Expert knowledge of parallel/distributed file systems such as Lustre or IBM GPFS
  • Strong knowledge of networking and cluster-based distributed computing, including InfiniBand switch configuration
  • Experience deploying open-source and commercial HPC platforms
  • HPC cluster architecture design and configuration; AI cluster and Kubernetes cluster experience
  • Strong scripting skills — Bash, Python, Perl; Ansible for automation is a plus
  • Experience building dashboards (e.g., Flask-based) for cluster monitoring is a plus
  • Skilled in diagnosing and debugging complex HPC hardware/software issues, with proposed workarounds
  • Strong troubleshooting and root-cause analysis capabilities
Key Skill Areas at a Glance
  • Networking: InfiniBand, InfiniBand Switch, Networking & Solution Design
  • Automation/Scripting: Python, Bash, Perl, Ansible, Flask Dashboards
  • OS Administration: Rocky Linux, Ubuntu, CentOS, RHEL, SUSE, Windows Server
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. HPC ENGINEER
Sr. HPC ENGINEER

Cognizant • Hyderabad

On-site
INR 800,000 - 1,200,000
Cognizant Hiring For Sr. HPC Engineer
Cognizant Hiring For Sr. HPC Engineer

Cognizant • Pune District, Chennai District, Bengaluru

On-site
INR 3,000,000 - 6,000,000
HPC Senior System Integrator/System Administrator
HPC Senior System Integrator/System Administrator

GBB • Mumbai

On-site
INR 1,000,000 - 1,500,000
HighPerformance Computing ( HPC) Administrator
HighPerformance Computing ( HPC) Administrator

VIT • Chennai District

On-site
INR 900,000 - 1,300,000
HPC Admin
HPC Admin

5 Star Recruitment • Chennai District

On-site
INR 2,000,000 - 4,000,000
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]

SanDisk • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Linux System Administrator
Linux System Administrator

SISL Global • Chennai District

On-site
INR 800,000 - 1,200,000
HPC Engineer
HPC Engineer

Whiteblue • Chennai

On-site
INR 1,500,000 - 2,500,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 900,000
Senior HPC Engineer SME/Architect
Senior HPC Engineer SME/Architect

Tata Consultancy Services • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,800,000 - 3,000,000