INFRASTRUCTURE AND CLUSTER ENGINEER (TS/SCI CLEARANCE REQUIRED)

NorthHill Technology

Springfield (VA)

On-site

USD 130,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NorthHill Technology Resources is seeking an Infrastructure and Cluster Engineer to administer and optimize a dedicated customer compute cluster, ensuring high availability and security. The role requires TS/SCI clearance and will involve Linux system administration, OpenShift, and AI workload orchestration.

You will maintain storage, InfiniBand networking, and the Run:AI workload manager, aligning with federal security standards and accreditation requirements. Springfield, VA-based, direct-hire.

Qualifications

  • 5+ years of Linux systems administration and infrastructure management experience.
  • Experience with high-performance computing environments and batch/workload orchestration.
  • Familiarity with OpenShift or Kubernetes and AI workload workflows.

Responsibilities

  • Manage day-to-day operations of the customer compute cluster and Linux systems.
  • Configure and optimize workload management using Run:AI for AI/ML workloads.
  • Tune hardware, OS, and network settings for performance and efficiency.

Skills

Linux Admin
Infrastructure Mgmt
HPC
OpenShift
Kubernetes
Run:AI
SLURM
Python
Bash
InfiniBand
GPU Networking

Tools

OpenShift
Kubernetes
Run:AI
SLURM
Python
Bash
InfiniBand

Job description

INFRASTRUCTURE AND CLUSTER ENGINEER (TS/SCI CLEARANCE REQUIRED)
  • Springfield, VA
  • Information Technology

NorthHill Technology Resources has a need for anInfrastructure and ClusterEngineerto support a Federal Program inSpringfield, VA.This is a direct-hire role with our client, a fast-growing Federal Integrator.An active TS/SCI Clearance is required, will sponsor for CI Polygraph within first year.

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration:Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management:Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization:Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management:Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration:Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance:Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Basic Qualifications:

  • Clearance:Active TS/SCI with the ability to obtain CI Poly.
  • Experience:5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
    • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
    • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
    • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
    • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus:Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Preferred Qualifications:

  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

Location: Springfield, VA

US Citizenship Required

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000
Infra & Cluster Engineer — TS/SCI Clearance
Infra & Cluster Engineer — TS/SCI Clearance

NorthHill Technology • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
CLOUD ENGINEER (TS/SCI CLEARANCE REQUIRED)
CLOUD ENGINEER (TS/SCI CLEARANCE REQUIRED)

Select Search Associates LLC • Fairfax (VA)

On-site
USD 120,000 - 160,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Systems Administrator - HPC Linux (TS/SCI Clearance Required)
Systems Administrator - HPC Linux (TS/SCI Clearance Required)

North Point Technology • Herndon (VA)

On-site
USD 120,000 - 150,000
Systems Administrator - HPC Linux (TS/SCI Clearance Required)
Systems Administrator - HPC Linux (TS/SCI Clearance Required)

North Point Technology • King of Prussia (PA)

On-site
USD 100,000 - 130,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
Systems Administrator - HPC Linux (TS/SCI Clearance Required)
Systems Administrator - HPC Linux (TS/SCI Clearance Required)

North Point Technology • Springfield (MO)

On-site
USD 90,000 - 130,000