Senior HPC DevOps Engineer

Peraton

United States

On-site

USD 146,000 - 234,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Peraton Labs seeks a poly cleared Senior HPC DevOps Engineer to own automation for an existing HPC/AI compute cluster (Linux) on-site near College Park, MD. You will codify repeatable ops with Ansible and drive state enforcement via an enterprise automation platform.

Responsibilities include drift detection, node onboarding, patch automation, logging, incident management, and runbooks. On-site work and strong security posture are required.

Qualifications

  • 12+ years of experience in relevant field and advanced degree combinations as stated in the ad.
  • 7+ years in Linux systems/SRE/DevOps with HPC or large-scale compute experience.
  • 3+ years building and operating Ansible automation at scale (roles/collections, idempotency, inventories, secrets).
  • Strong Linux hardening and compliance fundamentals (SELinux/AppArmor, SSH key automation).
  • Experience with clustered compute environments (HPC, large Linux farms).
  • Hands-on experience with container tooling and image lifecycle/versioning.
  • Active TS/SCI security clearance with polygraph.

Responsibilities

  • Automation ownership: manage automation workflows, inventories, credentials, RBAC, execution environments, and promotion across environments.
  • Enforce desired-state and drift detection across cluster services; implement alerting and reconcile runtime vs configured state.
  • Compute node onboarding (Bare-metal/VM): automated OS install/config, security baselines, scheduler/shared storage enrollment, hardware readiness checks.
  • Patch & vulnerability automation: rolling maintenance, image scanning, and lifecycle integration.
  • Logging & observability: emit auditable logs and integrate with metrics/alerting for incident response.
  • Incident/problem management: automate responses to common incidents using runbooks and out-of-band management.
  • Docs-as-code: versioned runbooks and operator guidance on documentation platform.

Skills

SRE/DevOps
Ansible automation
Linux hardening
Incident response
Documentation practices
Git workflows

Education

BS in computer science, IT, or related technical field
MS in computer science or related field
PhD in related field

Tools

Ansible
Git
Container tooling
CI/CD pipelines

Job description

Responsibilities

Peraton Labs is seeking a poly cleared Senior HPC DevOps Engineer to own the operations and automation lifecycle for an existing HPC/AI compute cluster (Linux). You will work closely with Peraton team members, as well as directly with our Maryland-based customer, in a fast-paced environment at a customer site. In this role you will codify repeatable operations in Ansible and drive execution through an enterprise automation controller to enforce desired state, detect drift, accelerate node onboarding, and streamline incident response via runbook automation integrated with monitoring and ITSM.

This position requires full-time on-site work at a customer site near College Park, MD.

Key responsibilities may include
  • Automation ownership: Own and manage automation workflows, including job templates, inventories, credentials, RBAC configurations, execution environments, and promotion across environments.
  • Desired-state and drift detection: Enforce desired state across cluster services via code-driven configuration; implement drift detection and alert on deviations; reconcile runtime state vs configured state.
  • Compute node onboarding (Bare-metal/VM): Build and maintain an automated node bootstrap workflow that installs/configures the OS, applies security and performance baselines, enrolls nodes into the scheduler and shared storage ecosystem, validates hardware and service readiness (CPU, network, accelerator, storage mounts), and reports pass/fail results.
  • Patch & vulnerability response: Implement rolling maintenance and patch automation to meet defined vulnerability response SLAs. Maintain version-controlled container build definitions and integrate image scanning into the build/release lifecycle.
  • Logging & observability: Ensure automation and operational workflows emit auditable logs to centralized analytics and integrate with metrics/alerting to enable reliable incident response, proactive detection, and safe auto-remediation.
  • Incident/problem management: Automate responses to common incidents (hung nodes, storage performance alarms, image vulnerabilities, hardware failures) leveraging out-of-band hardware management interfaces and standardized runbooks.
  • Docs-as-code: Keep runbooks and operational documentation versioned alongside automation and publish operator guidance to the orgs documentation platform.

*This position may be eligible for an increased sign-on bonus. Eligibility, bonus amount, and applicable terms and conditions will be discussed during the recruiting process

#MDFSP

#PLABS26

Qualifications

Required qualifications

  • 12+ years of experience and a BS in computer science, IT, or related technical field, MS and 10 years of experience, or a Ph.D. with 8 years of experience. Four years of additional experience is required in lieu of a Bachelors' degree for a total of 16 years of experience.
  • 7+ years in Linux systems / SRE / DevOps, including production cluster operations in an HPC or large-scale compute environment.
  • 3+ years of experience building and operating Ansible automation at scale (roles/collections, idempotency, inventories, secrets).
  • Strong Linux hardening & compliance fundamentals (SELinux/AppArmor, SSH key automation, baseline config management).
  • Demonstrated experience operating or automating clustered compute environments (HPC, large Linux farms, or similar).
  • Hands-on experience with container tooling in Linux environments, including image lifecycle/versioning.
  • Familiarity with incident response and runbook-driven operations; ability to automate common remediations.
  • Strong Git workflow and documentation practices.
  • Must hold at least one active/current technical certification from the following-
    • Systems engineering (e.g., INCOSE)
    • Information security (e.g., CISSP)
    • Networking (e.g., CCNA)
    • System Administration (e.g., RHCE, MCSE)
    • Virtualization (e.g., VCP)
    • IT systems management (e.g., ITIL)
    • Project management (e.g., PMP, Agile)
  • Active TS/SCI security clearance with a current polygraph is required

Preferred qualifications

  • Bare-metal provisioning experience (PXE/iPXE, Kickstart/Preseed, Foreman/MAAS) and hardware OOB management.
  • CI/testing for automation and promotion pipelines for playbooks
  • Experience with tuned performance profiles, HPC performance troubleshooting, and GPU node health validation.

#MDPM

#MDFSP

Peraton Overview

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world's leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can't be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we're keeping people around the world safe and secure.

Target Salary Range

$146,000 - $234,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual's experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

EEO

EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

All

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

External Job Posting Title Senior HPC DevOps Engineer
External Job Posting Title Senior HPC DevOps Engineer

Peraton • College Park (MD)

Hybrid
USD 146,000 - 234,000
Senior HPC DevOps Engineer
Senior HPC DevOps Engineer

Peraton • College Park (MD)

On-site
USD 146,000 - 234,000
Competitive salary
Potential for overtime
Discretionary bonus
Senior Infrastructure Software Engineer
Senior Infrastructure Software Engineer

Peraton • United States

On-site
USD 112,000 - 179,000
Sign-on bonus eligibility
Senior Infrastructure Software Engineer
Senior Infrastructure Software Engineer

Peraton • College Park (MD)

On-site
USD 112,000 - 179,000
Sign-on bonus
Senior Tech Lead - Cyber Systems Engineering
Senior Tech Lead - Cyber Systems Engineering

Peraton • College Park (MD)

On-site
USD 176,000 - 282,000
External Job Posting Title Senior Infrastructure Software Engineer
External Job Posting Title Senior Infrastructure Software Engineer

Peraton • College Park (MD)

On-site
USD 112,000 - 179,000
Linux Systems Engineer
Linux Systems Engineer

Peraton • United States

On-site
USD 86,000 - 138,000
Linux Systems Administrator (SDN / Cisco ACI)
Linux Systems Administrator (SDN / Cisco ACI)

Peraton • United States

On-site
USD 112,000 - 179,000
Technical Program Director / Lead Systems Engineer
Technical Program Director / Lead Systems Engineer

Peraton • College Park (MD)

On-site
USD 190,000 - 304,000
Sign-on bonus eligibility
Senior Tech Lead - Cyber Systems Engineering
Senior Tech Lead - Cyber Systems Engineering

Peraton • United States

On-site
USD 176,000 - 282,000
Sign-on bonus