Model Test and Measurement Engineer

DEFCON AI, Inc.

United States

Remote

USD 150,000 - 190,000

Full time

42 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Fully remote role
Equity package
Health insurance for you and family
Unlimited PTO
Flexible work environment
14 weeks of fully-paid parental leave

Job summary

DEFCON AI, Inc. is seeking an experienced Model Test & Measurement Engineer for an AI-enabled platform in a secure government cloud environment.

You will own the evaluation framework, build labeled ground truth datasets, and validate scoring and generative AI capabilities to provide trusted, explainable evidence for deployment decisions. You will work closely with data scientists and AI engineers while maintaining independence from model building and threshold approvals.

Qualifications

  • 5+ years with model validation as a named responsibility, not a side task.
  • Direct experience with ground-truth construction and sampling design.
  • Strong Python and SQL; comfort working independently from the teams whose models you evaluate.
  • Elevated personnel security requirements apply to portions of this work and are discussed during screening.

Responsibilities

  • Construct labeled ground truth for model evaluation.
  • Design and run sampled audits.
  • Gate releases against model version; maintain version inventory, evaluation records, and rollback triggers.
  • Run drift and override review.
  • Produce human-oversight and fairness / disparate-effect evidence.
  • Maintain independence from the build roles: produce evaluation evidence, but do not build the models, approve thresholds, or authorize releases against that evidence.
  • Measure workflow improvement: review time, throughput, backlog movement, and override and rework rates.
  • Design evaluation-phase QC: sample selection that does not mix populations, the unit of review, and fair comparison when methods or searched sources differ.
  • Define the measurement events other teams must emit, and confirm the review workspace emits workflow events from first use.

Skills

5+ years model validation
Ground-truth construction
Python
SQL

Job description

RESILIENCE IN THE FACE OF DISRUPTION. DEFCON AI is an insights company that leverages artificial intelligence, mathematical optimization, data analytics, and software engineering for resilient optimization of complex systems.
In today’s dynamically changing world, DEFCON AI’s technology aligns outcomes with operational goals, better decision making, and empowers customers to anticipate assess, and mitigate the impacts of disruptions.

Be the independent voice that keeps the whole program honest — your evaluation is the standard everyone else is held to.

About the Role

You’ll join the analytics and AI engineering team behind a system that genuinely matters: an AI-assisted platform that brings together records from dozens of disparate data sources, resolves them to the correct individual, highlights what analysts should review first, and provides transparent, explainable recommendations that users can trust. Operating within a secure government cloud environment, the platform tackles complex challenges in AI, data integration, and decision support where quality, trust, and accountability are mission-critical.

As a Model Test & Measurement Engineer, you’ll own the evaluation framework that helps ensure those systems perform as intended. You’ll build and maintain labeled ground truth datasets, design statistically sound audit and sampling methodologies, measure model performance across releases, and create the evidence packages that support deployment decisions. You’ll independently validate both scoring and generative AI capabilities, helping the team understand not only whether a model works, but how confidently its outputs can be trusted.

This is a role with genuine influence. Your assessments will inform release decisions, drive improvement efforts, and provide the objective evidence customers rely on when evaluating system performance. You’ll work closely with data scientists, AI engineers, and technical leadership while maintaining the independence needed to provide clear, defensible evaluations. You do not build the models you validate, and you do not approve thresholds or release authorization against your own evidence.

If you’re energized by measurement, validation, and making complex AI systems more trustworthy, this is an opportunity to have an outsized impact on both the technology and the mission it supports.

This is a fully remote role with occasional travel to DEFCON AI headquarters, customer sites, and partner facilities as needed.

Key Responsibilities
  • Construct labeled ground truth for model evaluation
  • Design and run sampled audits
  • Gate releases against model version; maintain version inventory, evaluation records, and rollback triggers
  • Run drift and override review
  • Produce human-oversight and fairness / disparate-effect evidence
  • Maintain independence from the build roles: produce evaluation evidence, but do not build the models, approve thresholds, or authorize releases against that evidence
  • Measure workflow improvement: review time, throughput, backlog movement, and override and rework rates
  • Design evaluation-phase QC: sample selection that does not mix populations, the unit of review, and fair comparison when methods or searched sources differ
  • Define the measurement events other teams must emit, and confirm the review workspace emits workflow events from first use
Required Qualifications
  • 5+ years with model validation as a named responsibility, not a side task
  • Direct experience with ground-truth construction and sampling design
  • Strong Python and SQL; comfort working independently from the teams whose models you evaluate
  • Elevated personnel security requirements apply to portions of this work and are discussed during screening
Preferred Qualifications
  • Regulated-industry or government model-risk background
  • Experience with fairness / disparate-effect testing and human-oversight documentation
  • NIST AI RMF or comparable practice
What Success Looks Like
  • Evaluation evidence a customer can rely on, independent of the teams that build the models
  • Releases gated against a clear version and evaluation record
  • Workflow metrics (review time, throughput, backlog, override/rework rates) that give the program an honest read on whether it's working
What We Offer
  • A fully remote, results-based environment
  • Competitive salary, bonus, and equity package
  • 100% employer paid, comprehensive health insurance including medical, dental, and vision for you and your family
  • Unlimited PTO, with your manager's approval
  • Flexible work environment where you manage your work day
  • 14 weeks of fully-paid parental leave

Salary Range: $150,000–$190,000. This represents the typical salary range for this position based on experience, skills, and other factors.

We’re an Equal Opportunity Employer: You’ll receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or on the basis of disability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Model Evaluation & Validation Engineer
Remote Model Evaluation & Validation Engineer

DEFCON AI, Inc. • United States

Remote
USD 150,000 - 190,000
Fully remote role
Equity package
Health insurance for you and family
+3
QA Engineer - Clearance Required New Remote, USA
QA Engineer - Clearance Required New Remote, USA

DEFCON AI, Inc. • Northern (KY)

Remote
USD 135,000 - 170,000
Fully remote work environment
Health insurance for you and family
Unlimited PTO
+2
Data & ML Engineer
Data & ML Engineer

Red Cell Partners • United States

On-site
USD 150,000 - 200,000
Fully remote work
Competitive salary + equity
100% employer paid health insurance
+2
Technical Director, AI Decision Systems
Technical Director, AI Decision Systems

DEFCON AI, Inc. • Northern (KY)

Remote
USD 220,000 - 260,000
Fully remote environment
Competitive salary, bonus & equity
Health insurance for you and your 
f​a
+3
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
Software Engineer, Agent Platform
Software Engineer, Agent Platform

Retool Inc. • San Francisco (CA)

On-site
USD 164,000 - 306,000
Hybrid work location
Medical, dental, vision
401(k)
Forward Deployed Engineer - Language Models
Forward Deployed Engineer - Language Models

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including **e
Software Engineer, Agent Platform San Francisco, United States
Software Engineer, Agent Platform San Francisco, United States

Retool, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 164,000 - 306,000
Hybrid work location
Research Engineer, Evals
Research Engineer, Evals

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+6
Machine Learning Research Scientist, Evaluations
Machine Learning Research Scientist, Evaluations

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2