Senior HPC DevOps Engineer | TS/SCI w/ MD POLY Security Clearance required

Capstone Technology Partners

College Park (MD)

On-site

USD 222,000 - 257,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Four weeks paid time off
Eleven paid holidays
401k with employer contributions and 3
Annual performance bonuses
Medical, Dental, and Vision

Job summary

Capstone Technology Partners seeks a Senior HPC DevOps Engineer to own the automation lifecycle for an HPC/AI compute cluster (Linux). You will codify operations with Ansible, enforce desired state, and accelerate onboarding while collaborating with our Maryland client in a fast-paced environment.

The role requires TS/SCI with MD polygraph, 12+ years' experience or advanced degrees, strong Linux hardening, and hands-on container tooling.

Qualifications

  • 12+ years of experience and a BS in computer science, IT, or related technical field, MS and 10 years of experience, or a Ph.D. with 8 years of experience.
  • 4 years of additional experience is required in lieu of a Bachelors’ degree for total of 16 years.
  • 7+ years in Linux systems / SRE / DevOps, including production cluster operations in HPC or large-scale compute environment.
  • 3+ years of experience building and operating Ansible automation at scale.
  • Strong Linux hardening & compliance fundamentals.
  • Demonstrated experience operating or automating clustered compute environments.
  • Hands-on experience with container tooling in Linux environments.
  • Familiarity with incident response and runbook-driven operations; automate remediation.
  • Strong Git workflow and documentation practices.
  • Active TS/SCI w/ Polygraph clearance required.

Responsibilities

  • Enforce desired state across cluster services via code-driven configuration; implement drift detection and alert on deviations.
  • Compute node onboarding: automate OS deployment, security baselines, scheduler enrollment, and hardware readiness checks.
  • Patch & vulnerability automation to meet SLAs; maintain container image definitions and integrate image scanning.
  • Logging and observability: emit auditable logs and integrate with metrics/alerts for reliable incident response.
  • Incident and problem management: automate responses to common incidents using runbooks.
  • Docs-as-code: maintain runbooks and operator guidance in the documentation platform.

Skills

Linux
SRE
DevOps
Ansible
Git workflow
Container tooling
HPC
Incident response

Education

Bachelor's degree in Computer Science or related field
Master's degree in CS/IT
PhD in CS/related field

Tools

PXE/iPXE
Kickstart/Preseed
Foreman/MAAS
Ansible Automation at scale

Job description

Senior HPC DevOps Engineer | TS/SCI w/ MD POLY Security Clearance required
  • College Park, MD

LCAT: IE3

Salary Range: $222- 257K

Location: College Park, MD

Security Clearance Required: TS/SCI with MD polygraph

Job Description:

We are seeking a fully cleared Senior HPC DevOps Engineer to own the operations and automation lifecycle for an existing HPC/AI compute cluster (Linux). You will work closely with team members, as well as directly with our Maryland-based customer, in a fast-paced environment. In this role you will codify repeatable operations in Ansible and drive execution through an enterprise automation controller to enforce desired state, detect drift, accelerate node onboarding, and streamline incident response via runbook automation integrated with monitoring and ITSM.

About US:

Capstone Technology Partners delivers mission-ready cybersecurity, engineering, and AI-enhanced software solutions for federal defense clients. We combine deep technical expertise with a people-first culture to help modernize securely, accelerate capability delivery, and strengthen operational resilience. Our work supports DoW, with proven experience in ICAM, Zero Trust, cloud-native engineering, and secure Agile development.

Capstone is built on Collaboration, Culture, Connection, and Community. Values that guide every engagement and every mission we support. We deliver secure, scalable solutions designed to meet the evolving demands of our clients’ operations.

Key Responsibilities may include:

  • Desired-state and drift detection: Enforce desired state across cluster services via code-driven configuration; implement drift detection and alert on deviations; reconcile runtime state vs configured state.
  • Compute node onboarding (Bare-metal/VM): Build and maintain an automated node bootstrap workflow that installs/configures the OS, applies security and performance baselines, enrolls nodes into the scheduler and shared storage ecosystem, validates hardware and service readiness (CPU, network, accelerator, storage mounts), and reports pass/fail results.
  • Patch & vulnerability response: Implement rolling maintenance and patch automation to meet defined vulnerability response SLAs. Maintain version-controlled container build definitions and integrate image scanning into the build/release lifecycle.
  • Logging & observability: Ensure automation and operational workflows emit auditable logs to centralized analytics and integrate with metrics/alerting to enable reliable incident response, proactive detection, and safe auto-remediation.
  • Incident/problem management: Automate responses to common incidents (hung nodes, storage performance alarms, image vulnerabilities, hardware failures) leveraging out-of-band hardware management interfaces and standardized runbooks.
  • Docs-as-code: Keep runbooks and operational documentation versioned alongside automation and publish operator guidance to the orgs documentation platform.

Required qualifications

12+ years of experience and a BS in computer science, IT, or related technical field, MS and 10 years of experience, or a Ph.D. with 8 years of experience. Four years of additional experience is required in lieu of a Bachelors’ degree for a total of 16 years of experience.

7+ years in Linux systems / SRE / DevOps, including production cluster operations in an HPC or large-scale compute environment.

3+ years of experience building and operating Ansible automation at scale (roles/collections, idempotency, inventories, secrets).

Strong Linux hardening & compliance fundamentals (SELinux/AppArmor, SSH key automation, baseline config management).

Demonstrated experience operating or automating clustered compute environments (HPC, large Linux farms, or similar).

Hands-on experience with container tooling in Linux environments, including image lifecycle/versioning.

Familiarity with incident response and runbook-driven operations; ability to automate common remediations.

Strong Git workflow and documentation practices.

Must hold at least one active/current technical certification from the following-

  • Information security (e.g., CISSP)
  • Networking (e.g., CCNA)
  • System Administration (e.g., RHCE, MCSE)
  • IT systems management (e.g., ITIL)

This position requires an active/current TS/SCI w/ Polygraph.

Preferred qualifications

  • Bare-metal provisioning experience (PXE/iPXE, Kickstart/Preseed, Foreman/MAAS) and hardware OOB management.
  • CI/testing for automation and promotion pipelines for playbooks
  • Experience with tuned performance profiles, HPC performance troubleshooting, and GPU node health validation.
  • Experience generation operational documentation from repos (Confluence).

Capstone offers a very competitive benefits package because we care about our employees. Our offerings including:

  • Four (4) weeks paid time
  • Eleven Paid Holidays
  • 401k plan with 4% employer contributions and immediate vesting plus annual 3% contruibution
  • Annual Bonuses for performance
  • Medical, Dental, and Vision and insurance
  • Education and Training budget provided
  • Technical Certification and coursework budget for computer/program expenses
  • Amazon Prime Membership Reimbursement
  • Birthday Club
  • Capstone Swag
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC DevOPS Engineer | TS/SCI w/MD poly required
Senior HPC DevOPS Engineer | TS/SCI w/MD poly required

Power3 • College Park (MD)

On-site
USD 222,000 - 257,000
Four weeks paid time off
11 paid holidays
401k with employer contributions
+4
Senior HPC DevOps Engineer-TS/SCI w/ MD Poly|HPC Clusters
Senior HPC DevOps Engineer-TS/SCI w/ MD Poly|HPC Clusters

Capstone Technology Partners • College Park (MD)

On-site
USD 222,000 - 257,000
Four weeks paid time off
Eleven paid holidays
401k with employer contributions and 3
+2
Senior HPC DevOps Engineer
Senior HPC DevOps Engineer

Peraton • Maryland

On-site
USD 120,000 - 150,000
Senior HPC DevOps Engineer
Senior HPC DevOps Engineer

Peraton • College Park (MD)

On-site
USD 146,000 - 234,000
Competitive salary
Potential for overtime
Discretionary bonus
System Administrator SA1 | TS/SCI w/ MD POLY Security Clearance required
System Administrator SA1 | TS/SCI w/ MD POLY Security Clearance required

Capstone Technology Partners • Hanover (MD)

On-site
USD 141,000 - 163,000
Four weeks paid time off
Eleven paid holidays
401k with employer contributions
+7
JavaScript Developer Software Engineer SWE1| TS/SCI w/ MD POLY Security Clearance required
JavaScript Developer Software Engineer SWE1| TS/SCI w/ MD POLY Security Clearance required

Capstone Technology Partners • Hanover (MD)

On-site
USD 150,000 - 170,000
Four weeks paid time off
Eleven paid holidays
401k with 4% employer contributions
+7
Senior Software Engineer- JAVA, SPRING BOOT, REST| SWE4| TS/SCI w/ MD POLY Security Clearance required
Senior Software Engineer- JAVA, SPRING BOOT, REST| SWE4| TS/SCI w/ MD POLY Security Clearance required

Capstone Technology Partners • Hanover (MD)

On-site
USD 238,000 - 275,000
Paid time off
Paid holidays
401k match
+5
Systems Administrator SA1 | TS/SCI w/MD poly required
Systems Administrator SA1 | TS/SCI w/MD poly required

Power3 Solutions • Hanover (MD)

On-site
USD 141,000 - 163,000
Systems Administrator SA1 | TS/SCI w/MD poly required
Systems Administrator SA1 | TS/SCI w/MD poly required

Power3 • Annapolis (MD)

On-site
USD 141,000 - 163,000
Four weeks PTO
Eleven Holidays
401k plan with employer contributions
+4
Senior DevOps Software Engineer (TS/SCI with Polygraph)
Senior DevOps Software Engineer (TS/SCI with Polygraph)

Red Alpha • Maryland

On-site
USD 175,000 - 220,000
Up to 10% in 401k contributions
5 weeks of leave (25 days personal time off)
100% health, dental, and vision insurance premiums
+3