Linux System Administrator

SISL Global

Chennai District

On-site

INR 800,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions company in Chennai seeks an HPC Engineer to manage and optimize High Performance Computing (HPC) environments with SLURM-managed CPU and GPU clusters. Responsibilities include user support, troubleshooting, and maintaining best practices for cluster operations. The ideal candidate will have strong Linux skills and experience with HPC architecture. This role offers the opportunity to work on cutting-edge technology in a collaborative environment.

Qualifications

  • Experience managing HPC clusters with SLURM in production environments.
  • Good understanding of Linux (RHEL) administration.
  • Knowledge of parallel computing concepts and HPC architecture.

Responsibilities

  • Manage day-to-day operations of on-prem HPC clusters including CPU and GPU compute nodes.
  • Monitor cluster health, performance, and utilization.
  • Implement and maintain best practices for HPC operations.

Skills

Experience managing HPC clusters with SLURM
Good understanding of Linux administration
Knowledge of parallel computing concepts
Strong troubleshooting and diagnostic skills

Tools

SLURM
Linux (RHEL)
WekaFS
Scality RING

Job description

Job Description – HPC Engineer (HPC with SLURM, CPU & GPU Clusters)

Position Overview

We are seeking a skilled HPC Engineer to design, deploy, manage, and optimize our on premises High Performance Computing (HPC) environment, consisting of SLURM-managed CPU and GPU clusters. The ideal candidate will have a strong understanding of HPC architecture, Linux systems, job scheduling, and cluster operations. Experience with parallel file systems and enterprise storage solutions such as WekaFS or Scality is preferred but optional.

Key Responsibilities
  • HPC Infrastructure & Operations
    • Manage day to day operations of on prem HPC clusters including CPU and GPU compute nodes.
    • Monitor cluster health, performance, and utilization, ensuring high availability and efficiency.
    • Implement and maintain best practices for HPC operations, user management, and resource administration.
    • Troubleshoot cluster related issues including networking, node failures, job failures, and performance bottlenecks.
    • Support users in job submissions, resource usage, and HPC workflows.
    • Configure, install, and manage SLURM workload manager across multiple clusters.
    • Handle queue creation, partition configuration, node allocation, fair share policies, and job prioritization.
    • Perform SLURM upgrades, migrations, and service maintenance with hands on expertise.
    • Work with SLURM APIs and integrations to support automation and custom workflows.
    • Optimize scheduling policies for mixed CPU/GPU workloads.
    • Manage Linux-based compute nodes, head nodes, and administration servers.
    • Perform OS updates, package installations, security patching, and system tuning.
    • Knowledge of shell scripting (Bash/Python) for automation and HPC tooling workflows.
  • Parallel Computing & Cluster Architecture
    • Understanding of parallel computing concepts: MPI, OpenMP, distributed execution.
    • Familiarity with HPC building blocks: interconnect networks (InfiniBand/100G), storage tiers, resource managers, monitoring tools.
    • Ability to analyze and troubleshoot performance issues in parallel workloads.
  • Storage (Optional but Preferred)
    • WEKA (WekaFS) – Optional
      • Knowledge of parallel file systems and performance tuning.
      • Diagnose and resolve issues related to WekaFS with minimal downtime.
      • Provide guidance to internal teams on WekaFS usage and best practices.
      • Stay updated with Weka ecosystem advancements and propose improvements.
    • Scality – Optional
      • Troubleshoot and maintain Scality RING and ARTESCA environments.
      • Monitor, tune, and optimize Scality-based storage for high availability and reliability.
      • Create and maintain documentation for Scality configuration and SOPs.
      • Recommend performance improvements based on new Scality enhancements.
Qualifications & Skills
  • Mandatory Skills
    • Experience managing HPC clusters with SLURM in production environments.
    • Good understanding of Linux (RHEL) administration.
    • Knowledge of parallel computing concepts and HPC architecture.
    • Strong troubleshooting and diagnostic skills.
    • Ability to work in complex, multi-node distributed environments.
  • Preferred/Optional Skills
    • Experience with WekaFS, Scality RING, or other parallel/distributed file systems.
    • Exposure to GPU computing (CUDA, NVIDIA drivers, GPU scheduling).
    • Familiarity with monitoring tools (Grafana, Prometheus).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Senior System Integrator/System Administrator
HPC Senior System Integrator/System Administrator

GBB • Mumbai

On-site
INR 1,000,000 - 1,500,000
HPC Engineer
HPC Engineer

PlusWealth Group • Gurugram District

On-site
INR 900,000 - 1,300,000
Medical insurance
Meals at office
Generous paid time off
Senior HPC Engineer
Senior HPC Engineer

Netweb Technologies India Ltd. • Faridabad District

On-site
INR 1,500,000 - 2,100,000
HPC Admin
HPC Admin

SHI Solutions India Pvt. Ltd. • Maharashtra

On-site
INR 1,000,000 - 1,500,000
HPC Engineer
HPC Engineer

Clovertex • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Lead HPC Engineer
Lead HPC Engineer

Clovertex • Hyderabad

On-site
INR 2,000,000 - 3,000,000
HPC Engineer
HPC Engineer

Whiteblue • Chennai

On-site
INR 1,500,000 - 2,500,000
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]

SanDisk • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Senior HPC Engineer
Senior HPC Engineer

Binaire Private Limited • New Delhi

On-site
INR 1,500,000 - 2,500,000
Opportunity to influence hardware selection
Ownership of high-performance compute infrastructure
HPC Admin - L2
HPC Admin - L2

SHI Solutions India Pvt. Ltd. • Bengaluru

On-site
INR 600,000 - 900,000