Company: Precision Solutions | Client: Synergist (Supporting: NSA)
Position: Data Scientist
Salary: $130k - 285k + benefits (Highly Dependent on Experience / All Levels Considered)
Location: Ft. Meade, MD | Full-time Onsite | 5 days a week
Clearance: Active TS/SCI with NSA Full Scope Polygraph Clearance Required (Maryland Client)
Summary
Since 2012, our client has helped mission-critical government organizations and businesses face their most daunting technology challenges. Their team has been a trusted partner to many government agencies and is extremely familiar with a wide variety of systems, policies, and procedures. Our client is also a distinguished custom software development firm dedicated to delivering premium solutions tailored for businesses and governmental needs. They are home to top-tier technology professionals recognized as industry pioneers, comprehensive engineers, and reliable consultants. These experts are adept at clear communication, excel in resolving complex challenges where others may falter, and are skilled in actualizing an organization’s vision.
Responsibilities
- Design, implement, and maintain AI model benchmark suites that measure model performance against defined evaluation criteria.
- Curate evaluation datasets, develop scoring methods, and produce clear and repeatable results reporting.
- Develop benchmarks and test scenarios that evaluate agentic workflows and multi-step task performance.
- Measure tool use, task completion, reliability, and other performance indicators across agentic systems.
- Curate, validate, and version evaluation datasets and test scenarios.
- Develop adversarial and edge-case datasets designed to stress-test model and agent behavior.
- Produce scoring frameworks, statistical analyses, visualizations, and written findings.
- Translate raw evaluation results into clear and actionable information for technical and nontechnical stakeholders.
- Document dataset lineage, scoring rationale, benchmark assumptions, evaluation methodology, and evaluation outputs.
- Collaborate with Researchers and Developers to operationalize evaluation workflows into stable and repeatable tools.
Requirements
- Experience with Python, InspectAI, MongoDB, and Jupyter notebooks.
- Experience designing or supporting AI model evaluation frameworks, LLM benchmarking, and agentic workflow evaluation.
- Experience evaluating tool use, multi-step task performance, task completion, and agent reliability.
- Strong knowledge of statistical analysis, experimental design, data cleaning, data validation, and data visualization.
- Experience curating, validating, and versioning evaluation datasets and test scenarios.
- Experience documenting dataset lineage, traceability, benchmark assumptions, scoring rationale, and evaluation results.
- Experience working within cloud-based or containerized environments.
- Experience using Git-based version control.
- Familiarity with Jira, Confluence, or similar collaboration and documentation tools.
- Ability to operate effectively within airgapped or otherwise constrained environments.
Technologies and Skills
- InspectAI
- Python
- MongoDB
- Jupyter notebooks
- Data visualization tools
- Statistical analysis
- Data cleaning and validation
- LLM benchmarking
- Agentic workflow evaluation
- Experimental design(? check balanced)