HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed

Springfield (VA)

On-site

USD 120,000 - 150,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

INflow Federal is seeking an Infrastructure & Cluster Engineer to manage the dedicated customer compute cluster, ensuring high availability, security, and optimized performance for complex workloads. You will work with Linux systems, OpenShift, and AI orchestration tools to deliver reliable infrastructure.

The role emphasizes hands-on management of bare-metal servers, InfiniBand networks, and GPU-to-GPU topology within an HPC context, with opportunities to grow expertise in Run:AI and SLURM and

Qualifications

  • 5+ years Linux systems administration in HPC environments.
  • Experience with InfiniBand networking and GPU-to-GPU topologies.
  • Proficiency with OpenShift or Kubernetes.
  • Experience with automation scripts (Bash, Python).
  • Knowledge of security/compliance standards.

Responsibilities

  • Cluster Administration: Manage day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

Skills

Linux Admin
HPC Environments
InfiniBand Network
OpenShift/Kubernetes
Run:AI/SLURM
Bash/Python Scripting

Tools

OpenShift
Kubernetes
Run:AI
SLURM
Bash
Python

Job description

INflow Federal is seeking an Infrastructure & Cluster Engineer to manage the dedicated customer compute cluster, ensuring high availability, security, and optimized performance for complex workloads. You will work with Linux systems, OpenShift, and AI orchestration tools to deliver reliable infrastructure.

The role emphasizes hands-on management of bare-metal servers, InfiniBand networks, and GPU-to-GPU topology within an HPC context, with opportunities to grow expertise in Run:AI and SLURM and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure & AI Compute Cluster Engineer
HPC Infrastructure & AI Compute Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
Senior HPC Cluster & Infra Engineer for AI Workloads
Senior HPC Cluster & Infra Engineer for AI Workloads

INflow Federal • Town of Springfield (WI)

On-site
USD 140,000 - 185,000
Travel opportunities
DoD 8140 certification training access
Career growth & learning
Lead HPC Cluster Engineer for AI/ML & OpenShift
Lead HPC Cluster Engineer for AI/ML & OpenShift

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Cluster Engineer: Secure, High-Performance Compute
HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Deployment Lead for AI & InfiniBand Systems
HPC Deployment Lead for AI & InfiniBand Systems

NVIDIA • Austin (TX)

On-site
USD 216,000 - 397,000
Equity
Benefits
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
HPC Infrastructure & Cluster Engineer
HPC Infrastructure & Cluster Engineer

INflow Federal • Town of Springfield (WI)

On-site
USD 140,000 - 185,000
Travel opportunities
DoD 8140 certification training access
Career growth & learning
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra

Jobgether • Germany (OH)

On-site
USD 81,000 - 105,000
Career development opportunities
Flexible working arrangements
Collaborative engineering environment
+1