HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig

Springfield (VA)

Hybrid

USD 140,000 - 185,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

INflow Federal is seeking an Infrastructure & Cluster Engineer to manage a dedicated customer compute cluster, ensuring high availability, security, and optimized performance. You will administer Linux systems, monitor hardware, patch, and upgrade while coordinating AI workloads across the cluster.

The role requires 5+ years in Linux administration within HPC settings, experience with InfiniBand, Run:AI, SLURM, OpenShift or Kubernetes, and scripting in Bash/Python.

Qualifications

  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus:Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Responsibilities

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Skills

Linux systems administration
High-performance computing
Troubleshooting

Tools

Run:AI
SLURM
OpenShift
Kubernetes
Bash
Python
InfiniBand

Job description

INflow Federal is seeking an Infrastructure & Cluster Engineer to manage a dedicated customer compute cluster, ensuring high availability, security, and optimized performance. You will administer Linux systems, monitor hardware, patch, and upgrade while coordinating AI workloads across the cluster.

The role requires 5+ years in Linux administration within HPC settings, experience with InfiniBand, Run:AI, SLURM, OpenShift or Kubernetes, and scripting in Bash/Python.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Infrastructure & AI Compute Cluster Engineer
HPC Infrastructure & AI Compute Cluster Engineer

INflow • Springfield (VA)

On-site
USD 140,000 - 185,000
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed • Springfield (VA)

On-site
USD 120,000 - 150,000
Senior HPC Cluster & Infra Engineer for AI Workloads
Senior HPC Cluster & Infra Engineer for AI Workloads

INflow Federal • Town of Springfield (WI)

On-site
USD 140,000 - 185,000
Travel opportunities
DoD 8140 certification training access
Career growth & learning
Lead HPC Cluster Engineer for AI/ML & OpenShift
Lead HPC Cluster Engineer for AI/ML & OpenShift

Abile Group, Inc • Springfield (VA)

On-site
USD 130,000 - 180,000
HPC Cluster Engineer: Secure, High-Performance Compute
HPC Cluster Engineer: Secure, High-Performance Compute

D2 Technical Services • Springfield (VA)

On-site
USD 170,000 - 180,000
Health/Dental/Vision
401(k) match
PTO
HPC Cluster Engineer for AI Workloads | TS/SCI
HPC Cluster Engineer for AI Workloads | TS/SCI

Socket.dev • Springfield (VA)

On-site
USD 148,000 - 179,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • California (MO)

On-site
USD 140,000 - 190,000
Senior Linux Architect for Secure HPC & DoD Missions
Senior Linux Architect for Secure HPC & DoD Missions

Inflow NS • Huntsville (AL)

On-site
USD 120,000 - 180,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Global HPC Network Engineer for AI Infra
Global HPC Network Engineer for AI Infra

Together • San Francisco (CA)

On-site
USD 190,000 - 280,000
Startup equity
Health insurance
Competitive benefits