Data Scientist

Peraton

Red Bank (AL)

On-site

USD 146,000 - 234,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Peraton is seeking a Data Scientist to build and curate a knowledge base for a generative AI/ML research effort, and to develop data and knowledge extraction pipelines for machine learning models. You will extract, structure, and validate design knowledge into knowledge graphs and ontologies; manage metadata, annotation, and provenance of program datasets.

The ideal candidate is a strong applied data scientist with an interest in symbolic and structured representations of knowledge.

Qualifications

  • Bachelor’s Degree or higher in Computer Science, Statistics, Mathematics, Electrical Engineering, or a related technical field.
  • 5+ years of applied data science or data engineering experience with data pipelines and analyses that others depend on.
  • Strong Python proficiency with NumPy, pandas, SciPy, scikit-learn and experience with at least one deep learning framework (PyTorch preferred).
  • Hands-on experience with knowledge graphs, ontologies, or structured knowledge representation and querying/validating them (SPARQL, Cypher, SHACL, etc.).
  • Experience designing data schemas, metadata standards, and annotation workflows for large datasets, including provenance.

Responsibilities

  • Design, build, and populate the research team’s knowledge bases with provenance tracking.
  • Develop knowledge extraction pipelines (rule-based, statistical, ML) to convert unstructured sources into structured knowledge.
  • Engineer metadata, annotation schemas, and packaging for program datasets to ensure reproducibility.
  • Perform exploratory and inferential analysis on training data and results; build dashboards and reports for model researchers.
  • Coordinate with subcontractors on shared data and knowledge resources.

Skills

Python proficiency
Statistical analysis
Deep learning (PyTorch)
SQL
Documentation

Education

Bachelor’s Degree or higher in Computer Science, Statistics, Mathematics, Electrical Engineering, or related field
MS or PhD in related field

Tools

Neo4j
RDF/OWL
SPARQL
Cypher
SHACL
Airflow
Dagster
DVC

Job description

Responsibilities

We are seeking a Data Scientist to build and curate the knowledge base for a generative AI/ML research effort, and to develop the data and knowledge extraction pipelines for machine learning models. The role is responsible for extracting, structuring, and validating design knowledge into knowledge graphs and ontologies; for engineering the metadata, annotation, and provenance of program datasets; and for the analysis that turns results, laboratory measurements, and simulation output into actionable findings for the research team.

The ideal candidate is a strong applied data scientist with an interest in symbolic and structured representations of knowledge. Experience with formal methods and domain-specific languages (DSLs) is desired but not required.

Key Responsibilities
  • Design, build, and populate the research team’s effort knowledge bases, and maintain it under version control with provenance tracking
  • Develop knowledge extraction pipelines (rule-based, statistical, and ML-assisted) that convert unstructured and semi-structured sources into structured, queryable knowledge; establish quality metrics and validation procedures for extracted content
  • Engineer the metadata, annotation schema, and packaging for program datasets (synthetic, simulated, and real collections) so that datasets are reproducible, well documented, and deliverable on the program data sharing schedule
  • Perform exploratory and inferential analysis on training data, simulation output, and evaluation results; build dashboards and reports that show where generated waveforms succeed or fail against objectives, and feed findings back to AI model researchers and engineers
  • Contribute data and analysis sections to design reviews, monthly status reports, and dataset documentation; coordinate with academic subcontractors on shared data and knowledge resources
Qualifications

Required Qualifications:

  • Bachelor’s Degree or higher in Computer Science, Statistics, Mathematics, Electrical Engineering, or a related technical field
  • 5+ years of applied data science or data engineering experience (or MS with 3+ years) with a record of delivering data pipelines and analyses that other engineers and researchers depend on
  • Strong Python proficiency including the scientific stack (NumPy, pandas, SciPy, scikit-learn) and experience with at least one deep learning framework (PyTorch preferred)
  • Hands‑on experience with knowledge graphs, ontologies, or structured knowledge representation (RDF/OWL, property graphs such as Neo4j, or equivalent) and with querying and validating them (SPARQL, Cypher, SHACL, or similar)
  • Experience designing data schemas, metadata standards, and annotation workflows for large scientific or engineering datasets, including data versioning and provenance
  • Solid statistical foundations: experimental design, hypothesis testing, uncertainty quantification, and the ability to explain results to technical and non‑technical audiences
  • Experience with SQL and with at least one workflow or pipeline orchestration tool (Airflow, Prefect, Dagster, DVC, or equivalent)
  • Ability to produce clear documentation including data dictionaries, dataset cards, and analysis reports
  • US Citizenship

Desired Qualifications:

  • Experience with formal methods or formal verification, such as SMT solvers (Z3, cvc5), model checkers, property‑based testing (Hypothesis), or proof assistants, particularly applied to validating generated programs or signal processing pipelines
  • Experience designing or implementing domain‑specific languages: grammar design, parser generators (ANTLR, Lark, or similar), type systems, intermediate representations, or compiling DSL programs to executable code
  • Background in symbolic AI or neuro‑symbolic methods: logic programming, constraint solving, rule engines, or program synthesis
  • Familiarity with digital signal processing and communications fundamentals (modulation, filtering, coding, channel effects) or with GNU Radio and software‑defined radio data formats (I/Q sample handling, SigMF or similar metadata standards)
  • Prior work on IARPA, DARPA, or similar government research programs, including data sharing plans, privacy protection plans, and delivery of datasets to independent T&E teams
  • Experience building causal or probabilistic models from structured knowledge, or working alongside causal inference researchers
  • MS or PhD in Computer Science, Statistics, Electrical Engineering, or a related technical field
  • Willingness and ability to obtain Secret security clearance
Peraton Overview

Peraton is a next‑generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Target Salary Range

$146,000 - $234,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

EEO

EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Peraton • Red Bank (NJ)

On-site
USD 146,000 - 234,000
Data Scientist Engineer Level 2
Data Scientist Engineer Level 2

Peraton • Laurel (MD)

On-site
USD 176,000 - 282,000
25 days PTO
Professional development opportunities
Bonus plan
Senior Data Scientist
Senior Data Scientist

Socket.dev • Washington

On-site
USD 146,000 - 234,000
Data Science, Associate - Herndon, VA
Data Science, Associate - Herndon, VA

Peraton • Herndon (VA)

On-site
USD 66,000 - 106,000
Data Scientist
Data Scientist

Peraton • Bowie (MD)

On-site
USD 66,000 - 106,000
Data Scientist
Data Scientist

Peraton • Maryland

On-site
USD 66,000 - 106,000
Senior Data Scientist
Senior Data Scientist

Peraton • Washington

On-site
USD 146,000 - 234,000
External Job Posting Title Data Science, Associate - Herndon, VA
External Job Posting Title Data Science, Associate - Herndon, VA

Peraton • Herndon (VA), Northern (KY)

Hybrid
USD 66,000 - 106,000
Data/Operations Research Analyst
Data/Operations Research Analyst

Peraton • Kentucky

On-site
USD 104,000 - 166,000
Senior Data Scientist
Senior Data Scientist

Peraton • Fort Meade (MD)

On-site
USD 146,000 - 234,000
AWS/Azure certifications