HPC Infrastructure & Cluster Engineer

General Dynamics Information Technology

Springfield (VA)

On-site

USD 150,000 - 190,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

401K match
Health benefits
Internal mobility
Education & certifications
Career growth
Paid vacation

Job summary

General Dynamics Information Technology is seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of a dedicated customer compute cluster. You will ensure a highly available, secure, and optimized hardware foundation.

Responsibilities include Linux administration, hardware monitoring, Run:AI/SLURM job scheduling, OpenShift/Kubernetes, InfiniBand networks, and automation with Bash/Python. Must have TS/SCI and 5+ years in HPC environments.

Qualifications

  • 5+ years of Linux systems administration and infrastructure management in HPC environments.
  • Experience with bare-metal servers and enterprise storage.
  • Experience with AI orchestration and workload management.
  • Automation scripting with Bash or Python.

Responsibilities

  • Manage day-to-day operations of the customer compute cluster (Linux admin, patching, upgrades).
  • Configure and optimize workload management and orchestration platforms (Run:AI, SLURM).
  • Tune performance across hardware, OS, and network levels for efficiency.
  • Administer storage and high-speed networks; support InfiniBand GPU-to-GPU topology.]
  • Provision environments and containers using OpenShift or Kubernetes; ensure security and compliance.

Skills

Linux systems administration
High-performance computing
Troubleshooting

Tools

Run:AI
SLURM
OpenShift
Kubernetes
Bash
Python
InfiniBand

Job description

Position Summary

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract at GDIT. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:
  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Basic Qualifications:
  • Clearance: Active TS/SCI with the ability to obtain CI Poly.
  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
    • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
    • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
    • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
    • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
Preferred Qualifications:
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

Location: Springfield, VA

US Citizenship Required

GDIT IS YOUR PLACE:
  • 401K with company match
  • Comprehensive health and wellness packages
  • Internal mobility team dedicated to helping you own your career
  • Professional growth opportunities including paid education and certifications
  • Cutting-edge technology you can learn from
  • Rest and recharge with paid vacation and holidays

#RoverGSS

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
INFRASTRUCTURE AND CLUSTER ENGINEER (TS/SCI CLEARANCE REQUIRED)
INFRASTRUCTURE AND CLUSTER ENGINEER (TS/SCI CLEARANCE REQUIRED)

NorthHill Technology • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Engineering Team Lead
HPC Engineering Team Lead

General Dynamics Corporation • Rockville (MD), Northern (KY)

Hybrid
USD 132,000 - 178,000
Software Integration Engineer 4
Software Integration Engineer 4

Praxis Engineering • Maryland

On-site
USD 195,000 - 263,000
Growth opportunities
Internal mobility support
Competitive pay
+1
HPC Cluster Engineer: AI Workloads & Secure Infra Ops
HPC Cluster Engineer: AI Workloads & Secure Infra Ops

General Dynamics Information Technology • Springfield (VA)

On-site
USD 150,000 - 190,000
401K match
Health benefits
Internal mobility
+3
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Technical Lead
HPC Technical Lead

General Dynamics Information Technology • Silver Spring (MD)

Hybrid
USD 212,000 - 286,000
Sr. Linux Systems Engineer (TS/Q Clearance Required)
Sr. Linux Systems Engineer (TS/Q Clearance Required)

General Dynamics Information Technology • Town of Germantown (WI)

On-site
USD 143,000 - 190,000
Health plan options
401(k) company match
Paid time off
Operations DevOps, Team Lead - TS/SCI - Sign-on Bonus!!! - Relocation assistance
Operations DevOps, Team Lead - TS/SCI - Sign-on Bonus!!! - Relocation assistance

General Dynamics Information Technology • Springfield (VA)

On-site
USD 145,000 - 196,000
401K with company match
Paid time off
Medical plan options