ML Engineer (Evaluation and Experimentation)

Cynnovative

Arlington (VA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cynnovative is seeking a Machine Learning Engineer (Evaluation & Experimentation) in Arlington, VA. The ideal candidate will design evaluation pipelines, run large-scale experiments, and ensure data quality. This role requires U.S. Citizenship and an active TS/SCI security clearance. Candidates should have a B.S. in a relevant field, proficiency in Python, and experience with data processing and statistical methods. Collaborate across teams to support national security efforts.

Qualifications

  • Strong communication skills and ability to collaborate cross-functionally.
  • Experience building experiment, evaluation, or analytics pipelines.
  • Familiarity with experiment tracking tools such as MLflow.

Responsibilities

  • Design and implement evaluation pipelines for LLM experimentation.
  • Execute statistical analysis and testing over experimental results.
  • Collaborate with ML systems engineers to ensure data capture.

Skills

Python
Data processing
Statistical methods
Version control systems (Git)
Collaboration
Experiment tracking tools (MLflow)

Education

B.S. in Computer Science, Data Science, or related field
M.S. or Ph.D. preferred

Tools

Statistical analysis tools

Job description

At Cynnovative, we leverage machine learning, computer science, and software engineering to address high-impact problems in the cyber domain, specifically those which are critical to U.S. national security. We primarily extend fundamental research to invent, design, develop, and deploy prototype solutions that support persistent problems in this domain.

Job Overview

As a Machine Learning Engineer (Evaluation & Experimentation) at Cynnovative, you will build and maintain systems that run large-scale experiments and evaluate LLM outputs. This role is crucial to rapid, experiment-driven iteration on LLM systems in support of U.S. national security efforts.

NOTE: This role requires an active TS/SCI security clearance and is located on-site in Northern Virginia.

Responsibilities (May Include)

Design and implement evaluation pipelines for LLM experimentation

  • Implement and apply metrics over model outputs at scale
  • Build automated evaluation workflows across large experiment sets
  • Execute statistical analysis and testing over experimental results
  • Ensure consistency and comparability of results across runs, configurations, and datasets

Develop experiment tracking and logging specifications

  • Define schemas for capturing prompts, perturbations, outputs, and configurations
  • Specify and validate logging of token-level probabilities, scores, and derived metrics
  • Ensure experiment data is structured, complete, and queryable for downstream analysis

Build and maintain datasets and evaluation inputs

  • Curate prompt sets, perturbation strategies, and test cases provided by the research team
  • Maintain versioned datasets and experiment inputs
  • Enable rapid iteration on experiment configurations and evaluation coverage

Collaborate cross-functionally

  • Work closely with ML systems engineers to ensure correct data capture at scale
  • Provide feedback on experiment execution, data quality, and metric behavior
  • Support interpretation of experimental results through reliable measurement
Requirements (Must Have)
  • B.S. in Computer Science, Data Science, or related field (M.S. or Ph.D. preferred)
  • Strong communication skills and ability to collaborate cross-functionally
  • Proficiency in Python and data processing
  • Experience building experiment, evaluation, or analytics pipelines
  • Familiarity with experiment tracking tools (MLflow or similar)
  • Experience working with large-scale or batch data processing workflows
  • Understanding of statistical methods
  • Experience working with structured and semi-structured data
  • Experience with version control systems, particularly Git
  • U.S. Citizenship and active TS/SCI security clearance
Desired Skills (Nice To Have)
  • Familiarity with prompt sensitivity, perturbation analysis, or robustness testing
  • Prior experience in a research-to-product environment
  • Understanding of A/B testing and large-scale experimentation
  • Familiarity with cyber-related data, tools, and techniques
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer (LLM Systems)
ML Engineer (LLM Systems)

Cynnovative • Arlington (VA)

On-site
USD 120,000 - 150,000
Software Engineer (ML Visualization & Tooling)
Software Engineer (ML Visualization & Tooling)

Cynnovative • Arlington (VA)

On-site
USD 85,000 - 120,000
ML Engineer: Evaluation & Experimentation (TS/SCI, NOVA)
ML Engineer: Evaluation & Experimentation (TS/SCI, NOVA)

Cynnovative • Arlington (VA)

On-site
USD 100,000 - 130,000
Senior Machine Learning Engineer - Mission Innovation Lab
Senior Machine Learning Engineer - Mission Innovation Lab

Software Engineering Institute | Carnegie Mellon University • Pittsburgh

On-site
USD 100,000 - 140,000
ML Experiment Visualization & Tools Engineer (TS/SCI)
ML Experiment Visualization & Tools Engineer (TS/SCI)

Cynnovative • Arlington (VA)

On-site
USD 85,000 - 120,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 210,000
Senior ML Engineer, LLM Systems — TS/SCI Cleared
Senior ML Engineer, LLM Systems — TS/SCI Cleared

Cynnovative • Arlington (VA)

On-site
USD 120,000 - 150,000
Lead, Data Engineering & Analytics - TS/SCI Required
Lead, Data Engineering & Analytics - TS/SCI Required

LMI • Reston (VA)

On-site
USD 120,000 - 150,000
Modeling & Simulation Analyst/Developer
Modeling & Simulation Analyst/Developer

LMI Government Consulting • Tysons (VA)

Hybrid
USD 95,000 - 120,000
Modeling & Simulation Analyst/Developer
Modeling & Simulation Analyst/Developer

LMI • Tysons (VA)

Hybrid
USD 100,000 - 130,000