HPC Infrastructure & Cluster Engineer

Abile Group, Inc

Springfield (VA)

On-site

USD 130,000 - 180,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Abile Group, Inc. seeks an experienced HPC Infrastructure & Cluster Engineer to support an Intelligence Community customer on a long-term contract. The role covers design, deployment, and operations of user-facing and data center IT services across networks and security domains worldwide.

The candidate will manage Linux HPC clusters, optimize workloads with Run:AI and SLURM, and implement OpenShift/Kubernetes environments while maintaining compliance with federal security standards.

Qualifications

  • Bachelor's Degree in a related discipline or equivalent education/training/experience.
  • 5+ years Linux systems administration and infrastructure management in HPC environments.
  • TS/SCI clearance with ability to obtain CI poly.
  • DoD 8570 IAT Level II certifications such as Security+ CE, CCNA, SSCP, GSEC, GICSP, CySA+.

Responsibilities

  • Cluster administration of customer compute cluster including Linux OS, hardware monitoring, patching, and upgrades.
  • Configure and optimize workload management and orchestration platforms (Run:AI, SLURM).
  • Tune cluster performance across hardware, OS, and network for workload throughput.
  • Manage storage and high-speed networks; support InfiniBand GPU-to-GPU network implementation.
  • Provision environments and containers (OpenShift/Kubernetes) for model deployment.
  • Ensure infrastructure compliance with federal security standards and attestations.

Skills

Bare-metal servers
Enterprise storage
InfiniBand
Workload managers
Job schedulers
AI orchestration
Run:AI
SLURM
Container orchestration
OpenShift
Kubernetes
Bash scripting
Python scripting
Troubleshooting

Education

Bachelor's Degree in related discipline

Tools

OpenShift
Kubernetes

Job description

Overview

Abile Group has an exciting and challenging opportunity for a HPC Infrastructure & Cluster Engineer on a 10 year contract providing User Facing and Data Center Services supporting an Intelligence Community customer. All the personnel on the team will work together to support innovative design, engineering, procurement, implementation, operations, sustainment and disposal of user facing and data center information technology (IT) services on multiple networks and security domains, at multiple locations worldwide, to support the IC mission.

The right candidate will possess the belowskills and qualificationsand be ready to handle all responsibilities independently and professionally.

Responsibilities
  • Cluster Administration:Manages the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management:Configures, maintains, and optimizes workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization:Tunes cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management:Administers storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration:Partners with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance:Ensures all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Qualifications

Clearance Required: TS/SCI with ability to obtain a CI Poly.

Degree and Years of Experience: Bachelor's Degree in a related discipline, or the equivalent combination of education, professional training, or work/military experience.

  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.

Required Certifications:

  • Meet DoD 8570 IAT Level II requirements including one of the following: Security+ CE, CND, SSCP, GSEC, GICSP, CySA+, or CCNA.

Required Skills:

  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus:Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Desired Skills:

  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
About Abile Group, Inc.

Abile Group, founded in July 2004 to support the Intelligence Community and its contractors across Enterprise Analytics, IT & Systems Engineering, and Program & Project Management, merged with Valiant Solutions in January 2026 - an established provider of cybersecurity technologies and services for Federal Agencies since 2005. Together, this partnership creates a stronger, more integrated cybersecurity organization with expanded opportunities for employees, deeper technical collaboration, and a unified mission. With significant experience serving the Federal Government, we remain dedicated to our employees and clients and seek high-performing professionals who excel at providing guidance, developing solutions, and delivering implementation support that blends industry best practices with client expertise and Abile’s broad technical capabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer, Sr.
DevOps Engineer, Sr.

Abile Group, Inc • Springfield (MO), Northern (KY)

Hybrid
USD 130,000 - 170,000
Cloud Engineer
Cloud Engineer

Abile Group, LLC • Springfield (VA)

On-site
USD 140,000 - 200,000
IEC Engineer
IEC Engineer

Abile Group • Springfield (AL)

On-site
USD 110,000 - 160,000
DevOps Engineer, Sr.
DevOps Engineer, Sr.

Abile Group, Inc • Springfield (VA)

On-site
USD 150,000 - 190,000
Systems Administrator (SaaS)
Systems Administrator (SaaS)

Abile Group, Inc • Springfield (MO), Northern (KY)

Hybrid
USD 90,000 - 120,000
Software Developer, Senior
Software Developer, Senior

Abile Group, Inc • Springfield (VA)

On-site
USD 150,000 - 230,000
Cloud Engineer
Cloud Engineer

Abile Group, Inc • Springfield (MO), Northern (KY)

Hybrid
USD 130,000 - 190,000
Cloud Engineer
Cloud Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 140,000 - 190,000
Cloud Operations and Customer Support Engineer, Jr.
Cloud Operations and Customer Support Engineer, Jr.

Abile Group, LLC • Springfield (VA)

On-site
USD 90,000 - 120,000
Systems Administrator (SaaS)
Systems Administrator (SaaS)

Abile Group, Inc • Springfield (VA)

On-site
USD 85,000 - 105,000