HPC Infrastructure & Cluster Engineer

Socket.dev

Springfield (VA)

On-site

USD 148,000 - 179,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health/Dental/Vision
401(k)
Paid Time Off
Stipends
Referral Bonuses

Job summary

Bcore is seeking an HPC Infrastructure & Cluster Engineer in Springfield, VA to lead the administration and optimization of a dedicated customer compute cluster. You will ensure a highly available, secure hardware foundation and partner with teams to deploy workloads using OpenShift, Kubernetes, and AI orchestration tools.

The role requires Active TS clearance (SCI eligibility) with CI Poly eligibility and 5+ years in Linux systems administration within HPC environments.

Qualifications

  • Requires Bachelor's degree.
  • 5+ years of Linux systems administration and infrastructure management in HPC.

Responsibilities

  • Manage the administration, health, and performance of the customer compute cluster.
  • Configure and optimize workload management and AI orchestration platforms (Run:AI/SLURM).
  • Tune hardware, OS, and network for maximum compute efficiency and data throughput.
  • Administer storage and InfiniBand network infrastructure.
  • Provision environments and container platforms (OpenShift).
  • Ensure security and compliance with federal standards.

Skills

InfiniBand
Run:AI
SLURM
OpenShift
Kubernetes
Bash
Python
Linux admin
Troubleshooting
Hardware management

Education

Bachelor's degree

Tools

OpenShift
Kubernetes
Run:AI
SLURM
InfiniBand
Bash
Python

Job description

Overview
HPC Infrastructure & Cluster Engineer Springfield, VA

Active TS (SCI eligibility) clearance and eligibility to obtain a CI poly

At Bcore, our strength comes from how we deliver impact to the mission. Whether it’s architecting critical IT solutions, producing actionable intelligence, or developing cutting edge technology, we succeed because of the expertise, collaboration, and agility of our teams. Our Mission Services division combines enterprise IT, cloud solutions, DevSecOps, systems engineering, software development, and operational support. Bcore accelerates decisive advantage for warfighters and intelligence professionals by fusing human insight, rapid-fire engineering, precision-measured outcomes, and relentless grit into mission-ready solutions.

Do you want to join a team that is building tailored technical solutions to modernize our government’s mission and our client’s business? Do you have a desire to change how people work? Are you interested in helping to protect our nation’s cyber interests? Join our growing team as a HPC Infrastructure & Cluster Engineer, supporting the NGAcustomer mission.

Responsibilities
What you get to do every day:

You will manage the administration, health, and performance of the foundational compute environment. You will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Qualifications

Clearance Required: Active TS clearance (with SCI Eligibility) and eligibility to obtain CI Poly

Education/Experience:

  • Requires Bachelor's degree
  • 5+years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.

Required Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM)
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues

What is ideal?

  • Familiarity with parallel file systems and high-throughput storage architecture.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies
  • Intelligence Community Experience preferred
What you can expect from us
  • Recognizing great achievements do not go unnoticed by bcore through service anniversaries, spot awards, and employee referral bonuses
  • You’ll join a growing organization of passionate, top-shelf, IT engineering professionals with extensive experience in actively developing the technology revolution in the Intelligence community
  • The expected salary range within the Washington, DC metropolitan area is: $148,000 - $179,000. Final compensation is unique to each individual and will be determined based on factors such as experience, education, geographic location, and contractual requirements. This is not a guarantee.
  • Benefits include Health/Dental/Vision, 401(k), Paid Time Off, STD/LTD/Life Insurance/Voluntary Life Insurance, Stipends, Referral Bonuses, and more.

BCore is proud to be an equal opportunity workplace. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Engineer
Cloud Engineer

Socket.dev • Springfield (VA)

On-site
USD 122,000 - 147,000
Health/Dental/Vision benefits
401(k)
Paid Time Off
+2
Software Developer-mid
Software Developer-mid

Socket.dev • Springfield (VA)

On-site
USD 134,000 - 162,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Linux Systems Engineer with Security Clearance
Linux Systems Engineer with Security Clearance

B/CORE • Springfield (VA)

On-site
USD 113,000 - 139,000
Health/Dental/Vision
401(k)
Paid Time Off
+3
Senior Software Developer with Security Clearance
Senior Software Developer with Security Clearance

B/CORE • Springfield (VA)

On-site
USD 148,000 - 180,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Software Engineer
Software Engineer

Bridge Core (BCore) • Herndon (VA)

On-site
USD 120,000 - 140,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Software Engineer
Software Engineer

Bridge Core • Herndon (VA)

On-site
USD 120,000 - 140,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Resident Architect with Security Clearance
Resident Architect with Security Clearance

B/CORE • Gaithersburg (MD)

On-site
USD 132,000 - 220,000
Health/Dental/Vision
401(k)
Paid Time Off
+4
Software Engineer
Software Engineer

B/CORE • Herndon (VA)

On-site
USD 120,000 - 140,000
Health/Dental/Vision
401(k)
Paid Time Off
+3
Software Engineer
Software Engineer

Bcore • Herndon (VA)

Hybrid
USD 180,000 - 195,000
Health/Dental/Vision
401(k)
Paid Time Off
+2
Network Engineer
Network Engineer

Bcore • Herndon (VA)

Hybrid
USD 140,000 - 170,000
Health/Dental/Vision
401(k)
Paid Time Off
+3