Senior HPC & Infrastructure Engineer

Soni

Cherry Hill Township (NJ)

On-site

USD 135,000 - 155,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Soni in Cherry Hill, NJ is seeking a Senior HPC & Infrastructure Engineer to lead Linux-based enterprise infrastructure across hybrid cloud and on-premises environments. You will manage HPC clusters, automation, and security while collaborating with researchers and engineers to enable advanced analytics and AI workloads.

The role emphasizes design, deployment, and ongoing optimization of scalable pipelines, with extensive experience in Azure, HPC scheduling, and performance tuning.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline, or equivalent professional experience.
  • 10+ years of experience managing Linux-based enterprise infrastructure, including high-performance computing environments.
  • Advanced expertise in Linux administration, preferably within RHEL or Rocky Linux ecosystems.
  • Hands-on experience with Microsoft technologies, including Azure, Entra ID, Active Directory, and Windows Server platforms.
  • Strong knowledge of HPC scheduling and workload management solutions such as Altair PBS, Slurm, or IBM Spectrum LSF.
  • Experience deploying and maintaining complex software stacks, development environments, and containerized applications.
  • Proficiency in automation and scripting using Bash, Python, or comparable technologies.
  • Experience supporting virtualized environments, workstations, and GPU-enabled computing platforms.
  • Strong understanding of infrastructure security principles, access controls, and best practices for protecting sensitive workloads.
  • Ability to partner effectively with engineering, research analytics, and technical operations teams.
  • Excellent troubleshooting, communication, and documentation skills.

Responsibilities

  • Lead design, deployment, and ongoing support of enterprise infrastructure across hybrid cloud and on-premises environments.
  • Administer large-scale Linux platforms, including clustered computing environments for modeling, simulation, analytics, and research.
  • Manage software ecosystems by installing, configuring, and maintaining development tools, scientific applications, compilers, libraries, and dependencies.
  • Analyze system performance and implement improvements to maximize efficiency and resource utilization.
  • Configure and support resilient Linux environments with fault tolerance, redundancy, and failover for networking and storage.
  • Oversee HPC resource management platforms, including job scheduling, queue administration, workload prioritization, and policy governance.
  • Collaborate with engineers, researchers, and business users to translate compute, storage, and workflow requirements into infrastructure solutions.
  • Maintain systems supporting data preparation, processing pipelines, and workflow orchestration across compute environments.
  • Design, implement, and manage Microsoft Azure environments, including government cloud deployments, with focus on security, governance, and compliance.
  • Support infrastructure required for AI and machine learning initiatives, ensuring performance, scalability, and reliability.
  • Establish and maintain security controls for critical systems, including identity management, auditing, data protection, encryption, and regulatory compliance.
  • Investigate and resolve issues affecting compute nodes, storage, software environments, scheduling platforms, and user workloads.
  • Develop automation solutions to streamline platform administration, deployments, monitoring, and operational consistency.
  • Produce and maintain technical documentation, system standards, procedures, and knowledge base materials.

Skills

Linux administration
HPC environments
Automation scripting
Cloud technologies
Security best practices
Communication
Documentation
Azure cloud

Education

Bachelor's degree in CS/IT/Engineering

Tools

Altair PBS
Slurm
IBM Spectrum LSF
Terraform
Azure Bicep
Ansible

Job description

Location: Cherry Hill, NJ – 3-4 days/week on-site

Position Overview

Our client is seeking a highly skilled Senior HPC & Infrastructure Engineer to support and evolve enterprise technology platforms spanning both cloud and on-premises environments. This role will serve as a technical lead for Linux-based infrastructure and high-performance computing (HPC) resources that enable advanced engineering, scientific, and data-intensive workloads. The ideal candidate will bring deep expertise in Linux administration, cloud technologies, automation, security, and performance optimization while partnering closely with technical stakeholders to deliver scalable and reliable computing solutions.

Key Responsibilities
  • Lead the design, deployment, and ongoing support of enterprise infrastructure solutions across hybrid cloud and on-premises environments.
  • Administer large-scale Linux platforms, including clustered computing environments used for modeling, simulation, analytics, and research activities.
  • Manage software ecosystems by installing, configuring, and maintaining development tools, scientific applications, compilers, libraries, and supporting dependencies.
  • Analyze system performance and implement improvements to maximize efficiency, throughput, and resource utilization.
  • Configure and support resilient Linux environments, incorporating fault tolerance, redundancy, and failover capabilities for networking and storage systems.
  • Oversee HPC resource management platforms, including job scheduling, queue administration, workload prioritization, and policy governance.
  • Collaborate with engineers, researchers, and business users to understand compute, storage, and workflow requirements and translate them into effective infrastructure solutions.
  • Maintain systems supporting data preparation, processing pipelines, and workflow orchestration across compute environments.
  • Design, implement, and manage Microsoft Azure environments, including government cloud deployments, with a focus on security, governance, and compliance.
  • Support infrastructure required for AI and machine learning initiatives, ensuring performance, scalability, and operational reliability.
  • Establish and maintain security controls for critical systems, including identity management, auditing, data protection, encryption, and regulatory compliance practices.
  • Investigate and resolve issues affecting compute nodes, storage systems, software environments, scheduling platforms, and user workloads.
  • Develop automation solutions to streamline platform administration, deployments, monitoring, and operational consistency.
  • Produce and maintain technical documentation, system standards, operational procedures, and knowledge base materials.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline, or equivalent professional experience.
  • 10+ years of experience managing Linux-based enterprise infrastructure, including high-performance computing environments.
  • Advanced expertise in Linux administration, preferably within Red Hat Enterprise Linux (RHEL) or Rocky Linux ecosystems.
  • Hands-on experience with Microsoft technologies, including Azure, Entra ID, Active Directory, and Windows Server platforms.
  • Strong knowledge of HPC scheduling and workload management solutions such as Altair PBS, Slurm, or IBM Spectrum LSF.
  • Experience deploying and maintaining complex software stacks, development environments, and containerized applications.
  • Demonstrated proficiency in automation and scripting using Bash, Python, or comparable technologies.
  • Experience supporting virtualized environments, workstations, and GPU-enabled computing platforms.
  • Strong understanding of infrastructure security principles, access controls, and best practices for protecting sensitive or regulated workloads.
  • Ability to partner effectively with engineering, research analytics, and technical operations teams.
  • Excellent troubleshooting, communication, and documentation skills.
Preferred Qualifications
  • Experience with Infrastructure-as-Code technologies such as Terraform or Azure Bicep.
  • Familiarity with configuration management platforms including Ansible, Puppet, or similar automation tools.
  • Background supporting scientific computing, parallel processing applications, or computational engineering environments.
  • Understanding of cybersecurity and compliance frameworks such as NIST, CIS Controls, or equivalent standards.
  • Experience supporting cloud-native AI, advanced analytics, or data science platforms.
Compensation:

$135,000 to $155,000 annually

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Linux Infrastructure Engineer
Senior Linux Infrastructure Engineer

Soni • Cherry Hill Township (NJ)

On-site
USD 125,000 - 145,000
Senior HPC Systems Engineer
Senior HPC Systems Engineer

Autonomai Recruitment • Chicago (IL)

On-site
USD 110,000 - 170,000
Senior HPC & Cloud Infrastructure Lead
Senior HPC & Cloud Infrastructure Lead

Soni • Cherry Hill Township (NJ)

On-site
USD 135,000 - 155,000
Senior Systems Engineer
Senior Systems Engineer

Holtec International • Camden (NJ)

On-site
USD 135,000 - 155,000
Medical, dental, and vision coverage
Hybrid work opportunities
401(k) with company match
+5
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
HPC System Administrator
HPC System Administrator

Cybotic System • Savannah (GA)

On-site
USD 90,000 - 150,000
HPC Operations Engineer
HPC Operations Engineer

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
Senior HPC Specialist (3202-1) Denver, CO
Senior HPC Specialist (3202-1) Denver, CO

ESR Healthcare • Denver (CO)

On-site
USD 90,000 - 130,000
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
Senior HPC DevOPS Engineer | TS/SCI w/MD poly required
Senior HPC DevOPS Engineer | TS/SCI w/MD poly required

Power3 • College Park (MD)

On-site
USD 222,000 - 257,000
Four weeks paid time off
11 paid holidays
401k with employer contributions
+4