Software Engineer-Data Engineering, Machine Learning (ML)

AAMVA (American Association of Motor Vehicle Administrators)

Arlington (VA)

Hybrid

USD 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

AAMVA (American Association of Motor Vehicle Administrators) is looking for a skilled Machine Learning Data Engineer to develop and operationalize ML solutions using cloud infrastructure (Azure or AWS). This role involves the full model lifecycle, from dataset preparation to deployment and performance monitoring.

The ideal candidate will have a Bachelor’s degree in a quantitative field and 3–5 years of hands-on experience in data engineering and machine learning. Strong communication skills and the ability to work with cross-functional teams are essential.

Qualifications

  • 3–5 years of experience in data engineering, ML engineering, or applied analytics.
  • Proficiency in Python for data processing and ML model development.
  • Experience building data pipelines that handle batch and streaming workloads.

Responsibilities

  • Designing and building dataset preparation pipelines for ML.
  • Training, evaluating, and tuning ML models.
  • Deploying models to production on cloud infrastructure.

Skills

Data engineering
Machine learning
Python
SQL
Cloud platform experience (Azure or AWS)
Feature engineering
Data pipelines
Statistical analysis
Problem-solving
Communication skills

Education

Bachelor's degree in computer science, data science, or related field

Tools

Azure Synapse Analytics
Apache Spark
Power BI
Git

Job description

Position Summary

The IT Division is responsible for the development and operations of information systems for the State and Federal agencies doing business related to or using information from the administration of motor vehicles and driver licenses. The Machine Learning (ML) Data Engineer position has core responsibilities for the design, development, deployment, and operational support of machine learning solutions on cloud infrastructure. This includes the full model lifecycle — from data acquisition and dataset preparation through feature engineering, experimentation, model training, validation, production deployment, and ongoing monitoring. Current applications include anomaly detection across high-volume messaging networks, but the scope encompasses any ML capability that strengthens system reliability, operational intelligence, and data-driven decision making across AAMVA systems.

Essential Duties and Responsibilities

We are seeking a talented Data Engineer with machine learning experience to join our team. You will design, build, and operationalize ML solutions running on cloud infrastructure (Azure or AWS). You will work across the full model lifecycle: preparing datasets, engineering features, running experiments, deploying models to production, and operating them on cloud infrastructure.

Key Responsibilities
  • Designing and building dataset preparation pipelines — acquiring, cleaning, transforming, and versioning data for ML training and evaluation
  • Engineering features that extract meaningful signals from structured and semi‑structured data sources (time‑series patterns, statistical profiles, categorical encodings)
  • Running structured experimentation — testing multiple algorithms against defined scenarios, measuring performance, and documenting findings
  • Training, evaluating, and tuning ML models including regression, classification, clustering, anomaly detection, and ensemble methods
  • Deploying models to production on cloud infrastructure and building the pipelines that keep them running (re‑training, scoring, threshold management)
  • Monitoring model performance in production — tracking drift, false positive rates, and detection efficacy over time
  • Building and maintaining batch and streaming data pipelines using Synapse, Fabric, Spark, and Event Hubs that feed ML systems
  • Writing and optimizing analytical queries (SQL, KQL, PySpark) for data exploration, statistical profiling, and real‑time analysis
  • Creating validation frameworks — synthetic test data generation, backtesting against historical logs, and shadow‑mode evaluation
  • Building dashboards and visualizations that communicate model outputs to technical and non‑technical stakeholders
  • Collaborating with cross‑functional teams to identify ML opportunities and translate operational problems into data solutions; communicating findings, trade‑offs, and model behavior clearly to technical and non‑technical audiences across IT, operations, and leadership.
Direct Reports

None

Qualifications

Formal Education: Bachelor's degree in computer science, data science, statistics, mathematics, or related quantitative field. Equivalent work experience may be substituted.

Key Knowledge, Skills, and Abilities
  • 3–5 years of hands‑on experience in data engineering, ML engineering, or applied analytics.
  • Hands‑on cloud platform experience (Azure or AWS) building and deploying data or ML solutions on managed cloud services; specific platform less important than depth of experience.
  • Working knowledge of statistical foundations: distributions, variance, standard deviation, trend vs. seasonality, hypothesis testing, and how to apply them to real operational data.
  • Experience with the ML experiment‑to‑production cycle: dataset preparation, feature engineering, model training, evaluation, and deployment.
  • Proficiency in Python for data processing, statistical analysis, and ML model development.
  • Strong SQL skills with understanding of relational database fundamentals: data modeling, query optimization, indexing strategies, and how SQL Server infrastructure supports production workloads (T‑SQL, stored procedures, Availability Groups).
  • Experience building data pipelines that handle batch and streaming workloads.
  • Experience with version control systems (Git) and CI/CD practices.
  • Strong problem‑solving skills, attention to detail, and ability to work independently on ambiguous problems.
  • Strong written and verbal communication skills — able to explain technical findings to non‑technical stakeholders and engage productively across IT, operations, and leadership; comfort operating outside the ML silo and contributing to broader technology discussions.
Preferred Qualifications
  • Experience with time‑series analysis, anomaly detection, or statistical process control on operational data.
  • Familiarity with unsupervised and semi‑supervised techniques (isolation forest, clustering, ensemble methods).
  • Experience building and managing ML model lifecycle on Azure (MLflow, Fabric ML, Azure ML) or AWS (SageMaker, Glue, Step Functions).
  • Familiarity with KQL (Kusto Query Language) for time‑series decomposition, log analytics, or real‑time data exploration.
  • Knowledge of data modeling and dimensional modeling concepts.
  • Experience with synthetic test data generation and model validation frameworks.
  • Familiarity with operations and monitoring of mission‑critical data platforms.
Technical Stack
  • Core Technologies: Microsoft Fabric, Azure Synapse Analytics, Apache Spark, Delta Lake, Azure Event Hubs.
  • ML & Analytics: scikit‑learn, PySpark ML, statistical modeling, time‑series analysis, feature engineering, model validation.
  • Languages: Python, SQL, PySpark, KQL, C#.
  • Data Infrastructure: T‑SQL, Stored Procedures, SQL Server Availability Groups.
  • Azure Services: Azure Functions, Azure Data Factory, Azure Key Vault.
  • Optional: Databricks, Snowflake, Lakehouse Architecture, Azure OpenAI; AWS candidates: equivalent services (SageMaker, Glue, Kinesis, Redshift) are acceptable in place of Azure‑specific stack items.
  • Visualization: Power BI.
  • Development: Azure DevOps, CI/CD.
Disclaimer Statement

The preceding job description has been written to reflect management’s assignment of essential functions. It does not prescribe or restrict the tasks that may be assigned.

Equal Opportunity Employer

AAMVA is an Equal Opportunity Employer / Veterans / Disabled.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer – AI/Machine Learning
Lead Data Engineer – AI/Machine Learning

Core Specialty • Cincinnati (OH)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision, and life ins.
Disability insurance
401(k) company-match
+5
Cloud ML Engineer: Data Pipelines & Anomaly Detection
Cloud ML Engineer: Data Pipelines & Anomaly Detection

AAMVA (American Association of Motor Vehicle Administrators) • Arlington (VA)

Hybrid
USD 90,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Bollinger Shipyards, Inc. • Raceland (LA)

On-site
USD 140,000 - 210,000
AI/ML Engineer
AI/ML Engineer

7Th Sky Tech • Minneapolis (MN)

Hybrid
USD 120,000 - 180,000
Hybrid work model
AI/ML Engineer
AI/ML Engineer

7Th Sky Tech • Charlotte (NC)

Hybrid
USD 120,000 - 180,000
AI/ML Engineer
AI/ML Engineer

7Th Sky Tech • Berkeley (CA)

Hybrid
USD 140,000 - 190,000
AI/ML Data Platform Engineer (Senior) (Onsite)
AI/ML Data Platform Engineer (Senior) (Onsite)

Serigor Inc • Linthicum (MD)

On-site
USD 170,000 - 210,000
Machine Learning Engineer
Machine Learning Engineer

Virtualitics • United States

On-site
USD 100,000 - 130,000
Competitive salary/equity/bonus
Comprehensive benefits package including medical, dental, and vision
AI-ML Engineer
AI-ML Engineer

Analytica • Bethesda (MD)

On-site
USD 120,000 - 180,000
Bonuses
Health insurance
Training funds
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Cloudflare • United States

On-site
USD 130,000 - 160,000
Equal Opportunity Employer
Diversity and Inclusiveness Initiatives
Reasonable accommodations for applicants with disabilities