AI Data Engineer - Permanent - Full Time

Purple Hires Inc

Princeton (NJ)

Hybrid

USD 150,000 - 210,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Purple Hires Inc in Princeton/NYC hybrid is seeking an AI Data Engineer to build and optimize ML data pipelines, model tracking, and governance integration. This role focuses on AI/ML infrastructures, data lineage, and scalable data processing.

You will collaborate with data scientists to implement feature stores, model versioning, and real-time monitoring while ensuring data quality and secure handling of sensitive AI artifacts.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
  • 5–7 years of hands‑on data engineering experience with at least 2 years in AI/ML workloads.
  • Proficiency in Python and ML frameworks (PyTorch, TensorFlow, scikit‑learn).
  • Strong experience with Databricks, Apache Spark, and distributed ML workflows.
  • Deep understanding of ML lifecycle: training, deployment, monitoring.
  • Experience with feature engineering, data preprocessing, and ML data pipelines.
  • Knowledge of vector databases, embeddings, and similarity search for AI apps.
  • Proficiency in SQL for structured and unstructured data management.
  • Understanding of data governance, model governance, and AI ethics.

Responsibilities

  • Design and implement data pipelines for AI model metadata, training data lineage, and model performance tracking.
  • Build data infrastructure on Databricks leveraging Spark for large-scale processing.
  • Develop MCP servers and enable AI data distribution via MCP.
  • Create feature engineering pipelines and data preprocessing workflows for AI training and inference.
  • Implement model versioning, experiment tracking, and model registry integration (MLflow or equivalent).
  • Create automated workflows for AI agent discovery, classification, and inventory management.
  • Design and maintain knowledge graph structures for AI model relationships and data lineage.
  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
  • Develop data quality frameworks specific to AI training and validation datasets.
  • Collaborate with data scientists to optimize data access and feature store implementations.
  • Implement security and compliance controls for sensitive AI data and artifacts.
  • Document AI data architectures, schemas, and integration patterns.

Skills

Python
ML Frameworks
Databricks
Spark
SQL
AI Governance
Collaboration
Data Quality
Model Monitoring
Analytics

Education

Bachelor's or Master's in CS/DS/ML

Tools

LangChain
HuggingFace
Neo4j
Amazon Neptune
Azure ML
SageMaker
Vertex AI

Job description

Role- AI Data Engineer
Location- Princeton, NJ & NYC, NY (Hybrid)
Role Overview

The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.

Key Responsibilities
  • Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
  • Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
  • Develop MCP servers and enable AI data distribution via MCP.
  • Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
  • Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
  • Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
  • Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
  • Develop data quality frameworks specific to AI training datasets and validation data.
  • Collaborate with data scientists to optimize data access patterns and feature store implementations.
  • Implement security and compliance controls for sensitive AI training data and model artifacts.
  • Create comprehensive documentation for AI data architectures, schemas, and integration patterns.
Required Skills and Qualifications
  • Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
  • 5-7 years of hands‑on experience in data engineering, with at least 2 years focused on AI/ML workloads.
  • Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
  • Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
  • Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
  • Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
  • Knowledge of vector databases, embeddings, and similarity search for AI applications.
  • Proficiency in SQL for structured and unstructured data management.
  • Understanding of data governance, model governance, and AI ethics principles.
  • Strong analytical and problem‑solving capabilities with attention to data quality.
  • Excellent collaboration skills for working with data scientists, ML engineers, and architects.
Preferred/Nice-to-Have Skills
  • Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine‑tuning.
  • Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
  • Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
  • Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
  • Understanding of AI model explainability and interpretability techniques.
  • Experience with A/B testing frameworks for ML model evaluation.
  • Certification in Databricks, AWS, Azure, or GCP AI/ML services.
  • Publications or contributions to open‑source ML projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Data Engineer
AI Data Engineer

AIT Global inc. • Princeton (NJ)

Hybrid
USD 120,000 - 180,000
AI Data Engineer
AI Data Engineer

Smart IT Frame LLC • Princeton (NJ)

On-site
USD 120,000 - 180,000
Senior AI Data Engineer/ Architect
Senior AI Data Engineer/ Architect

Altimetrik • New York (NY)

Hybrid
USD 140,000 - 190,000
Artificial Intelligence Data Engineer
Artificial Intelligence Data Engineer

Green Key Resources • New York (NY)

On-site
USD 150,000 - 190,000
Data AI Engineer
Data AI Engineer

Compunnel, Inc. • Columbus (OH)

On-site
USD 100,000 - 130,000
AI Data Engineer
AI Data Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 90,000 - 120,000
Senior AI Data Engineer
Senior AI Data Engineer

Karsun Solutions • Herndon (VA)

On-site
USD 165,000 - 180,000
Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks • Dallas (TX)

On-site
USD 150,000 - 190,000
Lead AI Engineer
Lead AI Engineer

Anblicks • Richardson (TX)

On-site
USD 180,000 - 240,000
Lead AI Engineer - Bengaluru, INDIA
Lead AI Engineer - Bengaluru, INDIA

Vytwo • Dallas (TX)

Hybrid
USD 120,000 - 160,000