HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR)

Springfield (VA)

On-site

USD 180,000 - 200,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Arena Technical Resources, LLC (ATR) is seeking an HPC Infrastructure and Cluster Engineer to manage a dedicated customer compute cluster in Springfield, VA. The role focuses on highly available, secure, and optimized hardware foundations for complex workloads.

The engineer will oversee Linux administration, hardware monitoring, and patching; configure AI workloads with Run:AI and SLURM; and ensure infrastructure meets federal security standards and accreditation requirements.

Qualifications

  • Active TS/SCI clearance with CI Poly eligibility.
  • 5+ years of Linux administration and infrastructure management in HPC.
  • Experience with InfiniBand and enterprise storage systems.
  • OpenShift or Kubernetes experience with AI orchestration.
  • Scripting in Bash and Python for automation.

Responsibilities

  • Administer and manage the customer compute cluster end-to-end.
  • Configure and optimize workload management and AI orchestration platforms (Run:AI).
  • Tune hardware, OS, and network for performance and throughput.
  • Manage storage and high-speed networking; InfiniBand topology.
  • Collaborate with integration teams to provision environments and containers (OpenShift).
  • Ensure security and compliance with federal standards.

Skills

TS/SCI clearance with CI Poly
Linux systems administration
HPC experience
InfiniBand networking
OpenShift / Kubernetes
Run:AI / SLURM
Scripting: Bash / Python
Troubleshooting

Education

Bachelor's Degree in Computer Science

Tools

OpenShift
Kubernetes
SLURM
Run:AI
Python
Bash

Job description

Job Title: HPC Infrastructure and Cluster Engineer
Job Location: Springfield, VA
Compensation: $180,000 - $200,000
Eligibility/Clearance: Candidate must possess an active TS/SCI Clearance and the ability to obtain a CI Polygraph
Job Description:

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities
  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Skills/Qualifications:
Required:
  • Clearance: Active TS/SCI with the ability to obtain CI Poly.
  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
Technical Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
Desired:
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
Education:

Bachelors Degree in Computer Science or a related field

ATR is an Equal Opportunity Employer (EOE) who will provide equal employment opportunity to employees and applicants for employment without regard to race, ethnicity, religion, color, sex, pregnancy, national origin, age, veteran status, ancestry, sexual orientation, gender identity or expression, marital status, family structure, genetic information, or mental or physical disability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
Senior HPC DevOPS Engineer | TS/SCI w/MD poly required
Senior HPC DevOPS Engineer | TS/SCI w/MD poly required

Power3 • College Park (MD)

On-site
USD 222,000 - 257,000
Four weeks paid time off
11 paid holidays
401k with employer contributions
+4
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Linux Systems Administrator - Top Secret Clearance
HPC Linux Systems Administrator - Top Secret Clearance

Jobot • Vicksburg (MS)

On-site
USD 120,000 - 180,000
Competitive Pay DOE
Comprehensive Benefits Package
401k with a match
+2
Systems Engineer - TS/SCI
Systems Engineer - TS/SCI

Xcelerate Solutions • Bethesda (MD)

On-site
USD 130,000 - 180,000
HPC Software Engineer
HPC Software Engineer

Cornerstone Defense LLC • Colorado Springs (CO)

On-site
USD 140,000 - 230,000
Systems Engineer - TS/SCI
Systems Engineer - TS/SCI

Xcelerate-Solutions-5 • Bethesda (MD)

On-site
USD 120,000 - 160,000
Systems Engineer - TS/SCI
Systems Engineer - TS/SCI

VMD Corp • Bethesda (MD)

On-site
USD 110,000 - 150,000