HPC Infrastructure & Cluster Engineer

Abile Group, LLC

Springfield (VA)

On-site

USD 140,000 - 190,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Abile Group is seeking an HPC Infrastructure & Cluster Engineer on a 10-year contract to support a Intelligence Community customer across multiple networks and locations. You will manage Linux clusters, optimize hardware and networks, and orchestrate AI workloads with RunAI and SLURM in OpenShift/Kubernetes environments.

The role requires a Bachelor's degree and 5+ years' Linux experience, with DoD 8570 IAT Level II certification or equivalent, TS/SCI with CI poly eligibility.

Qualifications

  • Bachelor's degree or equivalent experience in a related discipline.
  • At least 5+ years in Linux systems administration and infra management for HPC environments.
  • DoD 8570 IAT Level II certification or equivalent (Security+ CE, CND, SSCP, GSEC, GICSP, CySA+, CCNA).

Responsibilities

  • Manage day-to-day operations of customer compute clusters, including Linux admin and system upgrades.
  • Configure and optimize workload management and AI orchestration platforms (RunAI, SLURM).
  • Tune performance across hardware, OS, and networking to maximize throughput.
  • Administer storage and high-speed networks, including InfiniBand GPU-to-GPU topology.
  • Provision environments and containers using OpenShift/Kubernetes; automate maintenance with scripts.
  • Ensure security, compliance, and accreditation of all infrastructure components.

Skills

Linux systems administration
RunAI/SLURM workload management
OpenShift / Kubernetes
Automation scripting (Bash, Python)
Troubleshooting hardware/network/OS

Education

Bachelor's Degree in related discipline

Tools

InfiniBand networking

Job description

Abile Group has an exciting and challenging opportunity for a HPC Infrastructure & Cluster Engineer on a 10 year contract providing User Facing and Data Center Services supporting an Intelligence Community customer. All the personnel on the team will work together to support innovative design, engineering, procurement, implementation, operations, sustainment and disposal of user facing and data center information technology (IT) services on multiple networks and security domains, at multiple locations worldwide, to support the IC mission.

  • Cluster Administration Manages the day-to‑day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management Configures, maintains, and optimizes workload management and orchestration platforms, utilizing the RunAI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization Tunes cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management Administers storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU‑to‑GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration Partners with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance Ensures all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Clearance Required

TS/SCI with ability to obtain a CI Poly.

Degree and Years of Experience

Bachelor's Degree in a related discipline, or the equivalent combination of education, professional training, or work/military experience.

  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
Required Certifications
  • Meet DoD 8570 IAT Level II requirements including one of the following Security+ CE, CND, SSCP, GSEC, GICSP, CySA+, or CCNA.
Required Skills
  • Expertise in managing bare‑metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., RunAI, SLURM).
  • Hands‑on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus Proven ability to diagnose and resolve complex hardware, network, and OS‑level issues.
Desired Skills
  • Familiarity with parallel file systems and high‑throughput storage architectures.
  • Prior experience engineering or managing high‑speed GPU‑to‑GPU communication topologies.

Abile Group, founded in July 2004 to support the Intelligence Community and its contractors across Enterprise Analytics, IT & Systems Engineering, and Program & Project Management, merged with Valiant Solutions in January 2026 – an established provider of cybersecurity technologies and services for Federal Agencies since 2005. Together, this partnership creates a stronger, more integrated cybersecurity organization with expanded opportunities for employees, deeper technical collaboration, and a unified mission. With significant experience serving the Federal Government, we remain dedicated to our employees and clients and seek high‑performing professionals who excel at providing guidance, developing solutions, and delivering implementation support that blends industry best practices with client expertise and Abile’s broad technical capabilities.

Abile is committed to hiring the most qualified and best fit person for the job – always has, always will. Anyone requiring reasonable accommodations should email careers@abilegroup.com with requested details. A member of the HR team will respond to your request within 2 business days.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
Cloud Engineer
Cloud Engineer

Abile Group, Inc • Springfield (MO), Northern (KY)

Hybrid
USD 130,000 - 190,000
Cloud Engineer
Cloud Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 140,000 - 190,000
Hybrid Cloud Platform Engineer
Hybrid Cloud Platform Engineer

Abile Group, LLC • Springfield (VA)

Hybrid
USD 120,000 - 160,000
Systems Administrator (Desktop Support)
Systems Administrator (Desktop Support)

Abile Group, Inc • Springfield (VA)

On-site
USD 95,000 - 135,000
Cloud Engineer
Cloud Engineer

Abile Group, LLC • Springfield (VA)

On-site
USD 140,000 - 200,000
Senior Cloud Operations Solution Developer
Senior Cloud Operations Solution Developer

Abile Group, LLC • Springfield (VA)

On-site
USD 150,000 - 190,000
Systems Administrator (Desktop Support)
Systems Administrator (Desktop Support)

Abile Group, LLC • Springfield (VA)

On-site
USD 85,000 - 120,000
Onsite role in Springfield
Systems Administrator Advisor
Systems Administrator Advisor

Abile Group, LLC • St. Louis (MO)

On-site
USD 110,000 - 150,000
Systems Administrator (SaaS)
Systems Administrator (SaaS)

Abile Group, Inc • Springfield (VA)

On-site
USD 85,000 - 105,000