Data & ML Engineer

CV in

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DEFCON AI is seeking a Data & ML Engineer to build resilient data and model infrastructure for an AI-driven decision-support platform operated in accredited environments.

You will own the full lifecycle of data and models, ensure traceability, and manage probabilistic record linkage in a self-hosted or secure cloud context with frequent collaboration across teams.

Qualifications

  • 5+ years of experience in data engineering, data architecture, applied ML, ML engineering, or production analytics.
  • Strong command of Python and SQL for large, messy, operational datasets.
  • Proven ability to deliver systems for sustained operational use with traceability and validation.

Responsibilities

  • Architect and maintain a graph of entities and relationships for matching and provenance tracking.
  • Design probabilistic matching techniques, including blocking, candidate generation, scoring, clustering, and thresholds.
  • Build deduplication and known-record suppression to ensure data integrity and source custody.
  • Develop relevance and priority models for large, imperfect datasets with careful feature engineering.
  • Own calibration and threshold design, and define meaningful score interpretations.
  • Implement embeddings, vector storage, and retrieval across a provenance-tracked evidence base.
  • Integrate language models via managed services with self-hosted alternatives within security boundaries.
  • Create secure data pipelines for ingestion, transformation, validation, and publishing with audit logging.

Skills

Python
SQL
Data engineering
ML engineering
Communication

Education

Bachelor's degree in a relevant field

Tools

Cloud platforms
Vector databases
Graph databases

Job description

Data & ML Engineer at DEFCON AI.

About the role

This role is dedicated to constructing resilient data and model infrastructure for an AI-driven decision-support platform designed for accredited, controlled operational environments. The hire owns the full lifecycle of data and models, ensuring traceability and reliability from raw ingestion to explainable output. They are responsible for building a robust data and model layer that handles low-signal information while addressing challenges where standard accuracy metrics are insufficient. The position requires a unique blend of data engineering, machine learning, and software development to manage probabilistic record linkage and scoring capabilities. The engineer will create systems that are understandable and defensible, focusing on detailed design and data lineage. This is a hands-on position to build new, hardened functionality from the ground up, leveraging a mature platform for source custody and retrieval. The role directly tackles high-stakes challenges where the cost of specific errors demands specialized solutions.

Key facts

Location: Remote, USA

What you'll do
  • Architect and maintain a graph of entities and relationships to enable sophisticated matching logic and provenance tracking.
  • Design and implement probabilistic matching techniques, including blocking strategies, candidate generation, pairwise scoring, clustering, and threshold policy.
  • Build robust deduplication and known-record suppression mechanics to ensure data integrity and source custody.
  • Establish clear interface and data-flow documentation to serve as a definitive reference for the broader engineering team.
  • Develop relevance and priority models that function effectively across large, imperfect datasets with careful feature engineering.
  • Own calibration and threshold design, defining what a score truly represents and creating meaningful score thresholds.
  • Design abstention policies that identify uncertain or high-risk cases for human review, balancing false positives against false negatives.
  • Implement embeddings, vector storage, and retrieval mechanisms over a large, provenance-tracked evidence base.
  • Integrate language models through approved managed services while maintaining a self-hosted or open-weight alternative within the security boundary.
  • Bind generated text directly to its source records and treat a lack of evidence as a valid response rather than forcing a conclusion.
  • Build secure pipelines for ingestion, transformation, validation, and publishing across structured, semi-structured, and unstructured sources.
  • Implement quality checks, schema validation, lineage capture, and audit logging as standard expectations.
  • Create mechanisms for source drift detection to catch degradation before it impacts analysis.
  • Generate statistically representative synthetic data to enable development when live data access is restricted.
  • Manage the entire model lifecycle, including packaging, serving, versioning, and rollback capabilities.
  • Support occasional travel, estimated at a maximum of 25% of the time, for collaboration at headquarters and client engagements.
  • Use modern tools while maintaining a high standard of engineering rigor and detailed documentation.
Requirements
  • Possess 5+ years of cumulative experience in data engineering, data architecture, applied machine learning, ML engineering, or production analytics engineering.
  • Demonstrate a strong command of Python and SQL with proven ability to work on large, messy, operational datasets.
  • Show a history of delivering systems for sustained operational use, not just exploratory analysis or prototyping.
  • Exhibit proficiency in translating complex technical concepts for stakeholders who rely on the outcomes.
  • Meet US Citizenship requirements as a non-negotiable condition of employment.
  • Maintain Active US Secret clearance eligibility, as this work is performed in a controlled government cloud environment.
  • Adept at using AI-assisted development tools while maintaining a critical eye for verification and validation.
  • Capable of designing systems that are not only effective but also understandable and defensible under audit.
  • Comfortable working with probabilistic record linkage concepts and the implications of imperfect data.
  • Experienced with data lineage capture and ensuring a clear line of sight to original source materials.
  • Knowledge of how to implement scoring, ranking, and decision logic within complex datasets.
  • Familiarity with embedding techniques, vector databases, and retrieval-augmented generation principles.
  • Understanding of pipeline construction for diverse data formats and the implementation of validation checks.
  • Aware of the importance of schema management and audit logging in regulated environments.
  • Able to create and utilize synthetic data to support development cycles when live data is unavailable.
  • Proficient in establishing threshold policies and managing model versioning for rollback scenarios.
Nice to have
  • Active Top Secret clearance is highly valued for this role.
  • Direct experience applying probabilistic matching to inconsistent identity data.
  • Knowledge of failure modes in record linkage and master data management.
  • Background in identity management practices and graph data modeling.
  • Familiarity with controlled government cloud environments and their specific constraints.
  • Experience with language model integration via managed services and self-hosted alternatives.
  • Understanding of abstention policies and their implementation in scoring models.
  • Exposure to synthetic data generation for development and testing purposes.
Practical notes

This is a fully remote position located in the United States, requiring occasional travel estimated at a maximum of 25% of the time. The role operates within a controlled government cloud environment, necessitating strict adherence to US Citizenship and Active US Secret clearance requirements. Travel is primarily to support collaboration at DEFCON AI headquarters and client sites or vendor partners when necessary. The position is not intended for entry-level candidates and requires a significant background in the specified technical domains. All work is expected to adhere to high standards of engineering rigor, documentation, and auditability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data & ML Engineer
Data & ML Engineer

Red Cell Partners • United States

On-site
USD 150,000 - 200,000
Fully remote work
Competitive salary + equity
100% employer paid health insurance
+2
Remote Data & ML Engineer for AI-Powered Decisions
Remote Data & ML Engineer for AI-Powered Decisions

Doist • United States

Remote
USD 150,000 - 200,000
Fully remote work
Competitive salary + equity
100% employer paid health insurance
+2
Remote Data & ML Engineer for AI Decision Systems
Remote Data & ML Engineer for AI Decision Systems

DEFCON AI • McLean (VA)

On-site
USD 150,000 - 200,000
Fully remote work
Salary, bonus, and equity
Employer-paid health insurance for you
+3
Remote Data & ML Engineer for Resilient Defense AI
Remote Data & ML Engineer for Resilient Defense AI

DEFCON AI, Inc. • Northern (KY)

Hybrid
USD 150,000 - 200,000
Fully remote work
Health insurance
Unlimited PTO
+1
Remote Data & ML Engineer for Secure AI Systems
Remote Data & ML Engineer for Secure AI Systems

Red Cell Partners • United States

On-site
USD 150,000 - 200,000
Remote work
Bonus
Equity
+4
Technical Talent Acquisition Partner
Technical Talent Acquisition Partner

DEFCON AI, Inc. • Northern (KY)

Hybrid
USD 100,000 - 125,000
Fully remote
Bonus & equity
Health insurance for you & family
+3
Senior Data Scientist
Senior Data Scientist

ETHOS - Talent & Advisory • Baltimore (MD)

On-site
USD 140,000 - 210,000
Engineer, Cyber Security II
Engineer, Cyber Security II

TALENT Software Services • Columbia (SC)

On-site
USD 120,000 - 160,000
Forward Deployed Senior Machine Learning / Applied AI Engineer – 25% travel across D.C. and are[...]
Forward Deployed Senior Machine Learning / Applied AI Engineer – 25% travel across D.C. and are[...]

Orbis Group • Maryland

Hybrid
USD 170,000 - 210,000
Data Engineer
Data Engineer

Agile Defense • Colorado Springs (CO)

On-site
USD 150,000 - 170,000