HPC Systems Engineer

EITR Technologies LLC

Annapolis (MD)

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

EITR Technologies is hiring a Linux Systems Engineer to sustain large HPC environments supporting national security missions in an on-site Annapolis Junction facility.

You will manage production Linux systems, run upgrades, evaluate new hardware in labs, and collaborate with government stakeholders and OEMs to improve tooling and processes. The role requires TS/SCI clearance with FSP, DoD 8570 IAT II readiness, and strong scripting capabilities.

Qualifications

  • Bachelor's degree or 8+ years of relevant experience in lieu of a degree.
  • Strong hands-on Linux administration experience (Red Hat, CentOS, and/or SUSE).
  • Experience supporting HPC clusters, parallel file systems (Lustre, GPFS), high-speed interconnects (InfiniBand, Slingshot), large-scale storage systems, or client/server infrastructure.
  • Experience providing Tier 1/Tier 2 operational support in a production environment.
  • DoD 8570 IAT Level II certification (e.g., Security+ CE, CCNA Security, GSEC) — current, or ability to obtain prior to start.
  • Active TS/SCI clearance with Full Scope Polygraph completed within the last 7 years.

Responsibilities

  • Administer, patch, and troubleshoot enterprise Linux systems (Red Hat, CentOS, SUSE) in production HPC environments.
  • Support large parallel file systems (Lustre, GPFS) and high-speed interconnects (InfiniBand, Slingshot) including performance and health troubleshooting.
  • Provide Tier 1/Tier 2 operational support: triaging tickets, diagnosing node and job failures, coordinating hardware break/fix with vendors.
  • Handle node life cycle work at scale: provisioning, imaging, firmware updates, and validation of new hardware.
  • Script and automate routine tasks (Bash, Python) to keep thousands of nodes consistent.
  • Maintain security compliance in an accredited environment (STIGs, vulnerability remediation).
  • Participate in maintenance windows and system upgrades, and document procedures and runbooks as you go.
  • Get hands-on with emerging compute, storage, and networking technology through lab evaluations with vendor partners.

Skills

Linux administration
HPC clusters
Lustre
GPFS
InfiniBand
Tier 1/2 support
Bash scripting
Python scripting
Security compliance
STIGs
IAT II
TS/SCI clearance
Vendor coordination
Lab evaluations
Ansible
Monitoring (Nagios/Prometheus/Grafana)

Education

Bachelor's degree
8+ years experience in lieu of degree

Tools

Red Hat
CentOS
SUSE
Lustre
GPFS
InfiniBand

Job description

Location: Annapolis Junction, MD (100% on-site)

Clearance: Active TS/SCI with Full Scope Polygraph (within the last 7 years)

About EITR Technologies

EITR Technologies was built by technologists who are still billable today. We know what it means to be on-site and on-contract in the government space, and we built the company we always wished we worked for.

We're small on purpose. You won't be five layers deep in an org chart; you'll work directly with company leadership and senior government stakeholders, and you'll have a real say in our technical direction, tooling, and culture. Good ideas here turn into actual company offerings. We offer strong pay, industry-leading benefits, and a culture with no strings attached: come to our happy hours and game nights if you want, skip them if you don't. We hire adults and treat them like it.

About the Role

We're hiring a Linux Systems Engineer to help operate and sustain large high performance computing environments supporting national security missions. The scale of compute, storage, and interconnect in these environments is unusual even by HPC standards, and the team regularly works problems without a documented answer.

You'll support the full life cycle of HPC clusters and their surrounding infrastructure: keeping production systems healthy, supporting upgrades and modernization, and helping evaluate new hardware in a lab environment before it reaches mission networks. The work involves close collaboration with government stakeholders, engineers across multiple disciplines, and the OEMs whose gear you'll be running.

What You'll Do
  • Administer, patch, and troubleshoot enterprise Linux systems (Red Hat, CentOS, SUSE) in production HPC environments
  • Support large parallel file systems (Lustre, GPFS) and high-speed interconnects (InfiniBand, Slingshot), including performance and health troubleshooting
  • Provide Tier 1/Tier 2 support: triaging tickets, diagnosing node and job failures, and coordinating hardware break/fix with vendors
  • Handle node life cycle work at scale: provisioning, imaging, firmware updates, and validation of new hardware
  • Script and automate routine tasks (Bash, Python) to keep thousands of nodes consistent
  • Maintain security compliance in an accredited environment (STIGs, vulnerability remediation)
  • Participate in maintenance windows and system upgrades, and document procedures and runbooks as you go
  • Get hands-on with emerging compute, storage, and networking technology through lab evaluations with our vendor partners
Required Qualifications
  • Bachelor's degree and 3+ years of experience as a Systems Engineer or Systems Administrator supporting enterprise IT environments, or 8+ years of relevant experience in lieu of a degree
  • Strong hands-on Linux administration experience (Red Hat, CentOS, and/or SUSE) — this is non-negotiable
  • Experience supporting one or more of the following: HPC clusters, parallel file systems (Lustre, GPFS), high-speed interconnects (InfiniBand, Slingshot), large-scale storage systems, or client/server infrastructure
  • Experience providing Tier 1/Tier 2 operational support in a production environment
  • DoD 8570 IAT Level II certification (e.g., Security+ CE, CCNA Security, GSEC) — current, or ability to obtain prior to start
  • Active TS/SCI clearance with Full Scope Polygraph completed within the last 7 years
Nice to Have
  • Experience with HPC workload managers/schedulers (e.g., Slurm, PBS)
  • Scripting and automation experience (Bash, Python, Ansible)
  • Familiarity with monitoring and metrics platforms (e.g., Nagios, Prometheus, Grafana) at scale
  • Experience with hardware life cycle support in large data center environments (racking, cabling standards, firmware management)
  • Familiarity with security hardening and compliance in accredited DoD/IC environments (STIGs, vulnerability remediation)
  • Experience working directly with OEM vendors on escalations, RMAs, and pre-release hardware evaluation
Growth

This role is a strong fit for someone a few years into their sysadmin career who wants to go deep on HPC. You'll be working alongside senior engineers who have run large mission programs, and at a company our size, the path from "solid engineer" to lead, architect, or customer-facing roles is short and based entirely on performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

On-Site HPC Systems Engineer (TS/SCI)
On-Site HPC Systems Engineer (TS/SCI)

EITR Technologies LLC • Annapolis (MD)

On-site
USD 120,000 - 180,000
Technical Expert/Functional Expert (HPC)
Technical Expert/Functional Expert (HPC)

Reflexive Concepts, LLC • Georgetown (MD)

On-site
USD 140,000 - 210,000
HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000
Infrastructure Engineer
Infrastructure Engineer

GIGATEC Engineering • Maryland

On-site
USD 120,000 - 160,000
100% Paid Healthcare
10% 401k in every paycheck
100% Fully Vested!!!
Software Integration Engineer II
Software Integration Engineer II

TAP Engineering LLC • Maryland

On-site
USD 110,000 - 128,000
PTO 15-25 days + 11 holidays
401(k) match and profit sharing
Free medical coverage
+7
Systems Administrator - HPC Linux (TS/SCI Clearance Required)
Systems Administrator - HPC Linux (TS/SCI Clearance Required)

North Point Technology • Herndon (VA)

On-site
USD 120,000 - 150,000
Systems Integration Engineer
Systems Integration Engineer

Talentify • Columbia (MD)

On-site
USD 140,000 - 190,000
Systems Administrator 2 - HPC Support
Systems Administrator 2 - HPC Support

TAP Engineering • Maryland

On-site
USD 83,000 - 128,000
Paid Time Off
401(k) Match
Medical Coverage
+8
Systems Administrator - HPC Linux (TS/SCI Clearance Required)
Systems Administrator - HPC Linux (TS/SCI Clearance Required)

North Point Technology • Springfield (MO)

On-site
USD 90,000 - 130,000
Systems Administrator - HPC Linux (TS/SCI Clearance Required)
Systems Administrator - HPC Linux (TS/SCI Clearance Required)

North Point Technology • King of Prussia (PA)

On-site
USD 100,000 - 130,000