HPC Cluster Engineer: AI Workloads & Secure Infra Ops

General Dynamics Information Technology

Springfield (VA)

On-site

USD 150,000 - 190,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

401K match
Health benefits
Internal mobility
Education & certifications
Career growth
Paid vacation

Job summary

General Dynamics Information Technology is seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of a dedicated customer compute cluster. You will ensure a highly available, secure, and optimized hardware foundation.

Responsibilities include Linux administration, hardware monitoring, Run:AI/SLURM job scheduling, OpenShift/Kubernetes, InfiniBand networks, and automation with Bash/Python. Must have TS/SCI and 5+ years in HPC environments.

Qualifications

  • 5+ years of Linux systems administration and infrastructure management in HPC environments.
  • Experience with bare-metal servers and enterprise storage.
  • Experience with AI orchestration and workload management.
  • Automation scripting with Bash or Python.

Responsibilities

  • Manage day-to-day operations of the customer compute cluster (Linux admin, patching, upgrades).
  • Configure and optimize workload management and orchestration platforms (Run:AI, SLURM).
  • Tune performance across hardware, OS, and network levels for efficiency.
  • Administer storage and high-speed networks; support InfiniBand GPU-to-GPU topology.]
  • Provision environments and containers using OpenShift or Kubernetes; ensure security and compliance.

Skills

Linux systems administration
High-performance computing
Troubleshooting

Tools

Run:AI
SLURM
OpenShift
Kubernetes
Bash
Python
InfiniBand

Job description

General Dynamics Information Technology is seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of a dedicated customer compute cluster. You will ensure a highly available, secure, and optimized hardware foundation.

Responsibilities include Linux administration, hardware monitoring, Run:AI/SLURM job scheduling, OpenShift/Kubernetes, InfiniBand networks, and automation with Bash/Python. Must have TS/SCI and 5+ years in HPC environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Cluster Engineer: Secure, High-Performance Compute
HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
HPC Infrastructure & AI Compute Cluster Engineer
HPC Infrastructure & AI Compute Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed • Springfield (VA)

On-site
USD 120,000 - 150,000
Senior HPC Cluster & Infra Engineer for AI Workloads
Senior HPC Cluster & Infra Engineer for AI Workloads

INflow Federal • Town of Springfield (WI)

On-site
USD 140,000 - 185,000
Travel opportunities
DoD 8140 certification training access
Career growth & learning
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

General Dynamics Information Technology • Springfield (VA)

On-site
USD 150,000 - 190,000
401K match
Health benefits
Internal mobility
+3
Lead HPC Cluster Engineer for AI/ML & OpenShift
Lead HPC Cluster Engineer for AI/ML & OpenShift

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000