Experience: 5+ Years
Esconet Technologies is looking for a highly motivated HPC System Integrator/Administrator with a strong passion for cluster administration, system integration, and validation of HPC clusters. In this role, you will influence the overall integration, delivery, and management of largely open-source HPC products and solutions — spanning Intel and AMD processors, NVIDIA GPUs, storage, InfiniBand, and Linux software.
A solid understanding of parallel processing (problem decomposition and work distribution), parallel programming (MPI, OpenMP), and computer architecture is essential for this role.
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, Computational Science, equivalent mathematical sciences, or a related field, with 3+ years of relevant work experience; or an equivalent combination of education, training, and experience
- 3+ years of experience with software development in Linux
- 3+ years of experience with HPC clusters and systems integration
- Ability to manage the AI stack up to the framework level
- Experience with GPU cluster capabilities and storage integration (LUSTRE/BeeGFS, cluster storage)
- Working knowledge of object-based storage, networking, InfiniBand, solution designing, and technical documentation
Key Responsibilities
- Install, configure, fine-tune, and troubleshoot multi-vendor, multi-site Linux HPC servers
- Build and deploy open-source software as well as vendor/partner software
- Diagnose and resolve system operational issues quickly and effectively
- Verify full operation of systems, including network, systems, and storage performance
- Configure scheduling and queuing systems
- Assist technical support teams with questions and issues encountered by customers
- Coordinate with vendors to resolve hardware and software problems
- Document system administration procedures for routine and complex tasks (wikis)
- Maintain and monitor the security of HPC systems and servers
- Design and configure HPC/Kubernetes clusters with NVIDIA GPU support, including MIG (Multi-Instance GPU) and AI Factory deployments
- Deploy and manage cluster file systems such as Lustre, including OSS (Object Storage Server) and MDS (Metadata Server) node configuration
- Build and maintain effective working relationships with coworkers, managers, and clients
- Travel as needed for on-site cluster installation or maintenance (limited)
Desired Skills & Experience
- Building, configuring, and administering Linux distributions — Rocky Linux, Ubuntu, CentOS, RHEL, and SUSE
- Expert knowledge of parallel/distributed file systems such as Lustre or IBM GPFS
- Strong knowledge of networking and cluster-based distributed computing, including InfiniBand switch configuration
- Experience deploying open-source and commercial HPC platforms
- HPC cluster architecture design and configuration; AI cluster and Kubernetes cluster experience
- Strong scripting skills — Bash, Python, Perl; Ansible for automation is a plus
- Experience building dashboards (e.g., Flask-based) for cluster monitoring is a plus
- Skilled in diagnosing and debugging complex HPC hardware/software issues, with proposed workarounds
- Strong troubleshooting and root-cause analysis capabilities
Key Skill Areas at a Glance
- Networking: InfiniBand, InfiniBand Switch, Networking & Solution Design
- Automation/Scripting: Python, Bash, Perl, Ansible, Flask Dashboards
- OS Administration: Rocky Linux, Ubuntu, CentOS, RHEL, SUSE, Windows Server