HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services

Springfield (VA)

On-site

USD 170,000 - 180,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health/Dental/Vision
401(k) match
PTO

Job summary

D2 Technical Services is seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of a dedicated customer compute cluster. You will ensure a highly available, secure, and optimized hardware foundation for complex AI/ML workloads.

Responsibilities include Linux system administration, patching, and upgrades; configuring workload managers like Run:AI and SLURM; optimizing hardware, storage, and InfiniBand networks; and deploying OpenShift containers with

Qualifications

  • 5+ years of experience in Linux systems administration and infrastructure management.

Responsibilities

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations

Skills

Linux administration
OpenShift
Kubernetes
Python
Bash
Run:AI
SLURM
HPC

Tools

InfiniBand
Run:AI
SLURM
OpenShift
Kubernetes

Job description

D2 Technical Services is seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of a dedicated customer compute cluster. You will ensure a highly available, secure, and optimized hardware foundation for complex AI/ML workloads.

Responsibilities include Linux system administration, patching, and upgrades; configuring workload managers like Run:AI and SLURM; optimizing hardware, storage, and InfiniBand networks; and deploying OpenShift containers with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
HPC Infrastructure & AI Compute Cluster Engineer
HPC Infrastructure & AI Compute Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed • Springfield (VA)

On-site
USD 120,000 - 150,000
Lead HPC Cluster Engineer for AI/ML & OpenShift
Lead HPC Cluster Engineer for AI/ML & OpenShift

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
Senior HPC Cluster & Infra Engineer for AI Workloads
Senior HPC Cluster & Infra Engineer for AI Workloads

INflow Federal • Town of Springfield (WI)

On-site
USD 140,000 - 185,000
Travel opportunities
DoD 8140 certification training access
Career growth & learning
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Senior HPC Infrastructure Engineer: Clusters & Cloud
Senior HPC Infrastructure Engineer: Clusters & Cloud

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000