Remote Model Evaluation & Validation Engineer

DEFCON AI, Inc.

United States

Remote

USD 150,000 - 190,000

Full time

38 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote role
Equity package
Health insurance for you and family
Unlimited PTO
Flexible work environment
14 weeks of fully-paid parental leave

Job summary

DEFCON AI, Inc. is seeking an experienced Model Test & Measurement Engineer for an AI-enabled platform in a secure government cloud environment.

You will own the evaluation framework, build labeled ground truth datasets, and validate scoring and generative AI capabilities to provide trusted, explainable evidence for deployment decisions. You will work closely with data scientists and AI engineers while maintaining independence from model building and threshold approvals.

Qualifications

  • 5+ years with model validation as a named responsibility, not a side task.
  • Direct experience with ground-truth construction and sampling design.
  • Strong Python and SQL; comfort working independently from the teams whose models you evaluate.
  • Elevated personnel security requirements apply to portions of this work and are discussed during screening.

Responsibilities

  • Construct labeled ground truth for model evaluation.
  • Design and run sampled audits.
  • Gate releases against model version; maintain version inventory, evaluation records, and rollback triggers.
  • Run drift and override review.
  • Produce human-oversight and fairness / disparate-effect evidence.
  • Maintain independence from the build roles: produce evaluation evidence, but do not build the models, approve thresholds, or authorize releases against that evidence.
  • Measure workflow improvement: review time, throughput, backlog movement, and override and rework rates.
  • Design evaluation-phase QC: sample selection that does not mix populations, the unit of review, and fair comparison when methods or searched sources differ.
  • Define the measurement events other teams must emit, and confirm the review workspace emits workflow events from first use.

Skills

5+ years model validation
Ground-truth construction
Python
SQL

Job description

DEFCON AI, Inc. is seeking an experienced Model Test & Measurement Engineer for an AI-enabled platform in a secure government cloud environment.

You will own the evaluation framework, build labeled ground truth datasets, and validate scoring and generative AI capabilities to provide trusted, explainable evidence for deployment decisions. You will work closely with data scientists and AI engineers while maintaining independence from model building and threshold approvals.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Test and Measurement Engineer
Model Test and Measurement Engineer

DEFCON AI, Inc. • United States

Remote
USD 150,000 - 190,000
Fully remote role
Equity package
Health insurance for you and family
+3
Remote Data & ML Engineer for AI Decision Systems
Remote Data & ML Engineer for AI Decision Systems

DEFCON AI • McLean (VA)

On-site
USD 150,000 - 200,000
Fully remote work
Salary, bonus, and equity
Employer-paid health insurance for you
+3
QA Engineer for AI-Driven Decision Systems (Remote)
QA Engineer for AI-Driven Decision Systems (Remote)

DEFCON AI • McLean (VA)

On-site
USD 135,000 - 170,000
Fully remote
Health insurance
Unlimited PTO
+2
Remote AI Eval Engineer: Ground Truth & Datasets
Remote AI Eval Engineer: Ground Truth & Datasets

Entrada Ventures • United States

Hybrid
USD 120,000 - 210,000
Medical, dental & vision insurance
Equity and meaningful early-stage"
Unlimited PTO
+2
AI Security Research Engineer: Evaluation & Defense
AI Security Research Engineer: Evaluation & Defense

General Analysis • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Remote QA Engineer: AI-Driven Testing & Release Gate Expert
Remote QA Engineer: AI-Driven Testing & Release Gate Expert

DEFCON AI, Inc. • Northern (KY)

Hybrid
USD 135,000 - 170,000
Fully remote work environment
Health insurance for you and family
Unlimited PTO
+2
AI Model Evaluation Engineer – Build Verifications
AI Model Evaluation Engineer – Build Verifications

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
MacBook Pro
Premium AI development tools
OpenAI Codex Max
Remote Data & ML Engineer for Secure AI Systems
Remote Data & ML Engineer for Secure AI Systems

Red Cell Partners • United States

On-site
USD 150,000 - 200,000
Remote work
Bonus
Equity
+4
AI Evaluation & Model Risk Lead
AI Evaluation & Model Risk Lead

BrickRed Systems • Bellevue (WA)

On-site
USD 190,000 - 230,000
Remote AI Evaluation Data Engineer — Ground Truth
Remote AI Evaluation Data Engineer — Ground Truth

Propheticsoftware • Northern (KY)

Hybrid
USD 120,000 - 190,000