HPC Infrastructure & Cluster Engineer

D2 Technical Services

Springfield (VA)

On-site

USD 170,000 - 180,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health/Dental/Vision
401(k) match
PTO

Job summary

D2 Technical Services is seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of a dedicated customer compute cluster. You will ensure a highly available, secure, and optimized hardware foundation for complex AI/ML workloads.

Responsibilities include Linux system administration, patching, and upgrades; configuring workload managers like Run:AI and SLURM; optimizing hardware, storage, and InfiniBand networks; and deploying OpenShift containers with

Qualifications

  • 5+ years of experience in Linux systems administration and infrastructure management.

Responsibilities

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations

Skills

Linux administration
OpenShift
Kubernetes
Python
Bash
Run:AI
SLURM
HPC

Tools

InfiniBand
Run:AI
SLURM
OpenShift
Kubernetes

Job description

**ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations

Basic Qualifications

  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g. Run:AI, SLURM)
  • Hands‑on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
  • Experience writing automation and configuration scripts (e.g. Bash, Python) to streamline cluster maintenance
  • Proven ability to diagnose and resolve complex hardware, network, and OS-level issues

Preferred Qualifications

  • Familiarity with parallel file systems and high-throughput storage architecture
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies

Additional Information

  • All your information will be kept confidential according to EEO guidelines.
  • Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $170-$180k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
  • Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!

D2 Technical Services is committed to a merit-based recruitment process and encourages applications from all qualified individuals. As a Veteran‑Owned Small Business, we particularly welcome applications from veterans who have the requisite skills and experience. Job applicants that are interested in one of our openings and may require a reasonable accommodation to participate in the job application or interview process, should contact us to request an accommodation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000
Linux Systems Engineer - TS/SCI
Linux Systems Engineer - TS/SCI

Plus3 IT Systems • Charlottesville (VA)

On-site
USD 100,000 - 175,000
Health benefits
401(k) matching
Parental leave
HPC Cluster Engineer: Secure, High-Performance Compute
HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
Linux Systems Engineer - TS/SCI
Linux Systems Engineer - TS/SCI

Plus3 IT • Charlottesville (VA)

Hybrid
USD 100,000 - 175,000
Linux Systems Engineer - TS/SCI
Linux Systems Engineer - TS/SCI

Jobless • Charlottesville (VA), Northern (KY)

Hybrid
USD 100,000 - 175,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
HPC Platform Engineer
HPC Platform Engineer

Shield Consulting Solutions, Inc. • Maryland

On-site
USD 205,000 - 215,000
25 days PTO
11 paid holidays
Employer-paid healthcare for employees
+1
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Linux Systems Engineer
Linux Systems Engineer

D2 Consulting • Springfield (VA)

On-site
USD 125,000 - 135,000
Health/Dental/Vision
401(k) match
Accrued PTO
+3
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000