ML Engineer (Evaluation and Experimentation)

Cynnovative

Arlington (VA)

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cynnovative is seeking a Machine Learning Engineer (Evaluation & Experimentation) in Arlington, VA. The ideal candidate will design evaluation pipelines, run large-scale experiments, and ensure data quality. This role requires U.S. Citizenship and an active TS/SCI security clearance. Candidates should have a B.S. in a relevant field, proficiency in Python, and experience with data processing and statistical methods. Collaborate across teams to support national security efforts.

Qualifications

  • Strong communication skills and ability to collaborate cross-functionally.
  • Experience building experiment, evaluation, or analytics pipelines.
  • Familiarity with experiment tracking tools such as MLflow.

Responsibilities

  • Design and implement evaluation pipelines for LLM experimentation.
  • Execute statistical analysis and testing over experimental results.
  • Collaborate with ML systems engineers to ensure data capture.

Skills

Python
Data processing
Statistical methods
Version control systems (Git)
Collaboration
Experiment tracking tools (MLflow)

Education

B.S. in Computer Science, Data Science, or related field
M.S. or Ph.D. preferred

Tools

Statistical analysis tools

Job description

At Cynnovative, we leverage machine learning, computer science, and software engineering to address high-impact problems in the cyber domain, specifically those which are critical to U.S. national security. We primarily extend fundamental research to invent, design, develop, and deploy prototype solutions that support persistent problems in this domain.

Job Overview

As a Machine Learning Engineer (Evaluation & Experimentation) at Cynnovative, you will build and maintain systems that run large-scale experiments and evaluate LLM outputs. This role is crucial to rapid, experiment-driven iteration on LLM systems in support of U.S. national security efforts.

NOTE: This role requires an active TS/SCI security clearance and is located on-site in Northern Virginia.

Responsibilities (May Include)

Design and implement evaluation pipelines for LLM experimentation

  • Implement and apply metrics over model outputs at scale
  • Build automated evaluation workflows across large experiment sets
  • Execute statistical analysis and testing over experimental results
  • Ensure consistency and comparability of results across runs, configurations, and datasets

Develop experiment tracking and logging specifications

  • Define schemas for capturing prompts, perturbations, outputs, and configurations
  • Specify and validate logging of token-level probabilities, scores, and derived metrics
  • Ensure experiment data is structured, complete, and queryable for downstream analysis

Build and maintain datasets and evaluation inputs

  • Curate prompt sets, perturbation strategies, and test cases provided by the research team
  • Maintain versioned datasets and experiment inputs
  • Enable rapid iteration on experiment configurations and evaluation coverage

Collaborate cross-functionally

  • Work closely with ML systems engineers to ensure correct data capture at scale
  • Provide feedback on experiment execution, data quality, and metric behavior
  • Support interpretation of experimental results through reliable measurement
Requirements (Must Have)
  • B.S. in Computer Science, Data Science, or related field (M.S. or Ph.D. preferred)
  • Strong communication skills and ability to collaborate cross-functionally
  • Proficiency in Python and data processing
  • Experience building experiment, evaluation, or analytics pipelines
  • Familiarity with experiment tracking tools (MLflow or similar)
  • Experience working with large-scale or batch data processing workflows
  • Understanding of statistical methods
  • Experience working with structured and semi-structured data
  • Experience with version control systems, particularly Git
  • U.S. Citizenship and active TS/SCI security clearance
Desired Skills (Nice To Have)
  • Familiarity with prompt sensitivity, perturbation analysis, or robustness testing
  • Prior experience in a research-to-product environment
  • Understanding of A/B testing and large-scale experimentation
  • Familiarity with cyber-related data, tools, and techniques
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer (LLM Systems)
ML Engineer (LLM Systems)

Cynnovative • Arlington (VA)

On-site
USD 120,000 - 150,000
Software Engineer (ML Visualization & Tooling)
Software Engineer (ML Visualization & Tooling)

Cynnovative • Arlington (VA)

On-site
USD 85,000 - 120,000
ML Engineer: Evaluation & Experimentation (TS/SCI, NOVA)
ML Engineer: Evaluation & Experimentation (TS/SCI, NOVA)

Cynnovative • Arlington (VA)

On-site
USD 100,000 - 130,000
ML Experiment Visualization & Tools Engineer (TS/SCI)
ML Experiment Visualization & Tools Engineer (TS/SCI)

Cynnovative • Arlington (VA)

On-site
USD 85,000 - 120,000
Senior ML Engineer, LLM Systems — TS/SCI Cleared
Senior ML Engineer, LLM Systems — TS/SCI Cleared

Cynnovative • Arlington (VA)

On-site
USD 120,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

Adelphi Data • McLean (VA)

Hybrid
USD 140,000 - 180,000
Principal LLM Application Engineer
Principal LLM Application Engineer

system_two • Palo Alto (CA)

Remote
USD 180,000 - 260,000
Medical, Dental, Vision
Equity in the company
M&S Software Developer
M&S Software Developer

LMI • Washington

Hybrid
USD 111,000 - 193,000
Forward Deployed Engineer - Clearance Required
Forward Deployed Engineer - Clearance Required

LMI Consulting, LLC • United States

On-site
USD 150,000 - 230,000
Lead AI Engineer
Lead AI Engineer

The Sherwin-Williams Company • Cleveland (OH)

On-site
USD 120,000 - 160,000