Lead HPC Cluster Engineer for AI/ML & OpenShift

Abile Group, Inc

Springfield (VA)

On-site

USD 130,000 - 180,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Abile Group, Inc. seeks an experienced HPC Infrastructure & Cluster Engineer to support an Intelligence Community customer on a long-term contract. The role covers design, deployment, and operations of user-facing and data center IT services across networks and security domains worldwide.

The candidate will manage Linux HPC clusters, optimize workloads with Run:AI and SLURM, and implement OpenShift/Kubernetes environments while maintaining compliance with federal security standards.

Qualifications

  • Bachelor's Degree in a related discipline or equivalent education/training/experience.
  • 5+ years Linux systems administration and infrastructure management in HPC environments.
  • TS/SCI clearance with ability to obtain CI poly.
  • DoD 8570 IAT Level II certifications such as Security+ CE, CCNA, SSCP, GSEC, GICSP, CySA+.

Responsibilities

  • Cluster administration of customer compute cluster including Linux OS, hardware monitoring, patching, and upgrades.
  • Configure and optimize workload management and orchestration platforms (Run:AI, SLURM).
  • Tune cluster performance across hardware, OS, and network for workload throughput.
  • Manage storage and high-speed networks; support InfiniBand GPU-to-GPU network implementation.
  • Provision environments and containers (OpenShift/Kubernetes) for model deployment.
  • Ensure infrastructure compliance with federal security standards and attestations.

Skills

Bare-metal servers
Enterprise storage
InfiniBand
Workload managers
Job schedulers
AI orchestration
Run:AI
SLURM
Container orchestration
OpenShift
Kubernetes
Bash scripting
Python scripting
Troubleshooting

Education

Bachelor's Degree in related discipline

Tools

OpenShift
Kubernetes

Job description

Abile Group, Inc. seeks an experienced HPC Infrastructure & Cluster Engineer to support an Intelligence Community customer on a long-term contract. The role covers design, deployment, and operations of user-facing and data center IT services across networks and security domains worldwide.

The candidate will manage Linux HPC clusters, optimize workloads with Run:AI and SLURM, and implement OpenShift/Kubernetes environments while maintaining compliance with federal security standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
HPC Infrastructure & AI Compute Cluster Engineer
HPC Infrastructure & AI Compute Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
Senior HPC Cluster & Infra Engineer for AI Workloads
Senior HPC Cluster & Infra Engineer for AI Workloads

INflow Federal • Town of Springfield (WI)

On-site
USD 140,000 - 185,000
Travel opportunities
DoD 8140 certification training access
Career growth & learning
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed • Springfield (VA)

On-site
USD 120,000 - 150,000
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
HPC Cluster Engineer: Secure, High-Performance Compute
HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package