Data Scientist with Security Clearance

Precision Solutions Ag

Fort Meade (MD)

On-site

USD 130,000 - 285,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Precision Solutions is seeking a Data Scientist to design and maintain AI model benchmark suites, curate evaluation datasets, and develop scoring methods for rigorous model assessment at Ft. Meade, MD.

The role emphasizes translating evaluation results for stakeholders and collaborating to deploy stable evaluation tools. Requirements include strong Python, InspectAI, MongoDB, and analytics skills, with experience in AI benchmarking and cloud/container environments.

Qualifications

  • Experience with Python and data analysis.
  • Experience designing or supporting AI model evaluation frameworks, LLM benchmarking, and agentic workflow evaluation.
  • Experience evaluating tool use, multi-step task performance, task completion, and agent reliability.
  • Strong knowledge of statistical analysis, experimental design, data cleaning, data validation, and data visualization.
  • Experience curating, validating, and versioning evaluation datasets and test scenarios.
  • Experience documenting dataset lineage, traceability, benchmark assumptions, scoring rationale, and evaluation results.
  • Experience working within cloud-based or containerized environments.
  • Experience using Git-based version control.
  • Familiarity with Jira, Confluence, or similar collaboration and documentation tools.
  • Ability to operate effectively within airgapped or otherwise constrained environments.

Responsibilities

  • Design, implement, and maintain AI model benchmark suites that measure model performance against defined evaluation criteria.
  • Curate evaluation datasets, develop scoring methods, and produce clear and repeatable results reporting.
  • Develop benchmarks and test scenarios that evaluate agentic workflows and multi-step task performance.
  • Measure tool use, task completion, reliability, and other performance indicators across agentic systems.
  • Curate, validate, and version evaluation datasets and test scenarios.
  • Develop adversarial and edge-case datasets designed to stress-test model and agent behavior.
  • Produce scoring frameworks, statistical analyses, visualizations, and written findings.
  • Translate raw evaluation results into clear and actionable information for technical and nontechnical stakeholders.
  • Document dataset lineage, scoring rationale, benchmark assumptions, evaluation methodology, and evaluation outputs.
  • Collaborate with Researchers and Developers to operationalize evaluation workflows into stable and repeatable tools.

Skills

Python
InspectAI
MongoDB
Jupyter notebooks
Data visualization
Statistical analysis
Data cleaning
LLM benchmarking
Agentic workflow evaluation
Experimental design

Tools

Git
Jira
Confluence

Job description

Company: Precision Solutions | Client: Synergist (Supporting: NSA)
Position: Data Scientist
Salary: $130k - 285k + benefits (Highly Dependent on Experience / All Levels Considered)
Location: Ft. Meade, MD | Full-time Onsite | 5 days a week
Clearance: Active TS/SCI with NSA Full Scope Polygraph Clearance Required (Maryland Client)

Summary

Since 2012, our client has helped mission-critical government organizations and businesses face their most daunting technology challenges. Their team has been a trusted partner to many government agencies and is extremely familiar with a wide variety of systems, policies, and procedures. Our client is also a distinguished custom software development firm dedicated to delivering premium solutions tailored for businesses and governmental needs. They are home to top-tier technology professionals recognized as industry pioneers, comprehensive engineers, and reliable consultants. These experts are adept at clear communication, excel in resolving complex challenges where others may falter, and are skilled in actualizing an organization’s vision.

Responsibilities
  • Design, implement, and maintain AI model benchmark suites that measure model performance against defined evaluation criteria.
  • Curate evaluation datasets, develop scoring methods, and produce clear and repeatable results reporting.
  • Develop benchmarks and test scenarios that evaluate agentic workflows and multi-step task performance.
  • Measure tool use, task completion, reliability, and other performance indicators across agentic systems.
  • Curate, validate, and version evaluation datasets and test scenarios.
  • Develop adversarial and edge-case datasets designed to stress-test model and agent behavior.
  • Produce scoring frameworks, statistical analyses, visualizations, and written findings.
  • Translate raw evaluation results into clear and actionable information for technical and nontechnical stakeholders.
  • Document dataset lineage, scoring rationale, benchmark assumptions, evaluation methodology, and evaluation outputs.
  • Collaborate with Researchers and Developers to operationalize evaluation workflows into stable and repeatable tools.
Requirements
  • Experience with Python, InspectAI, MongoDB, and Jupyter notebooks.
  • Experience designing or supporting AI model evaluation frameworks, LLM benchmarking, and agentic workflow evaluation.
  • Experience evaluating tool use, multi-step task performance, task completion, and agent reliability.
  • Strong knowledge of statistical analysis, experimental design, data cleaning, data validation, and data visualization.
  • Experience curating, validating, and versioning evaluation datasets and test scenarios.
  • Experience documenting dataset lineage, traceability, benchmark assumptions, scoring rationale, and evaluation results.
  • Experience working within cloud-based or containerized environments.
  • Experience using Git-based version control.
  • Familiarity with Jira, Confluence, or similar collaboration and documentation tools.
  • Ability to operate effectively within airgapped or otherwise constrained environments.
Technologies and Skills
  • InspectAI
  • Python
  • MongoDB
  • Jupyter notebooks
  • Data visualization tools
  • Statistical analysis
  • Data cleaning and validation
  • LLM benchmarking
  • Agentic workflow evaluation
  • Experimental design(? check balanced)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Precision Solutions, Inc. • Fort Meade (MD)

On-site
USD 100,000 - 130,000
Data Scientist
Data Scientist

Precision Solutions Ag • Fort Meade (MD)

On-site
USD 120,000 - 160,000
AI Software Engineer with Security Clearance
AI Software Engineer with Security Clearance

Precision Solutions Ag • Fort Meade (MD)

Hybrid
USD 130,000 - 290,000
AI Researcher
AI Researcher

Precision Solutions Ag • Fort Meade (MD)

On-site
USD 110,000 - 170,000
ME00654-Data Scientist 2
ME00654-Data Scientist 2

Momentum Engineering • Fort Meade (MD)

On-site
USD 105,000 - 140,000
11 paid holidays
3 weeks PTO
Company-sponsored group medical plan
+4
ME00654-Data Scientist 2
ME00654-Data Scientist 2

Momentum Engineering, Inc • Fort Meade (MD)

On-site
USD 105,000 - 140,000
11 paid holidays
3 weeks PTO
Medical plan
+4
ME00654-Data Scientist 2
ME00654-Data Scientist 2

Momentum Engineering, Inc. • Fort Meade (MD)

On-site
USD 105,000 - 140,000
3 weeks PTO
Holiday benefits
Company-sponsored health plan
+2
ME00648-Data Scientist 4
ME00648-Data Scientist 4

Momentum Engineering • Fort Meade (MD)

On-site
USD 170,000 - 220,000
11 paid holidays
3 weeks PTO
Medical plan
+4
Computer Scientist
Computer Scientist

Phase2 Technology • McLean (VA)

On-site
USD 99,000 - 225,000
Computer Scientist
Computer Scientist

Booz Allen Hamilton • Washington

On-site
USD 99,000 - 225,000