AI Data Engineer

Reuben Cooley Inc.

Jersey City (NJ)

Hybrid

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Reuben Cooley Inc. seeks an AI Data Engineer to design and optimize data pipelines for AI model metadata, training data lineage, and model performance metrics.

You will build data platforms on Databricks using Spark for large-scale processing and enable AI data distribution via MCP. You will implement model versioning and experiment tracking with MLflow, develop feature engineering workflows, and collaborate with data scientists to ensure secure, governance-aligned data access and quality.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Science, ML, or related field.
  • 5-7 years of hands-on data engineering, with at least 2 years in AI/ML workloads.
  • Proficient in Python and SQL; strong data governance and security awareness.
  • Experience with Databricks, Spark, and MLflow for ML data pipelines.

Responsibilities

  • Design and implement data pipelines for AI model metadata and training data lineage.
  • Build data infrastructure on Databricks leveraging Spark for large-scale processing.
  • Develop ML workflows, feature engineering, and data preprocessing for training/inference.
  • Implement model versioning, experiment tracking, and model registry integration.
  • Create automated workflows for AI agent discovery, classification, and inventory.
  • Design knowledge graphs representing AI model relationships and data lineage.
  • Build real-time data pipelines for model monitoring and drift detection.
  • Develop data quality frameworks for AI training and validation data.
  • Collaborate with data scientists to optimize feature stores and data access.
  • Implement security and compliance controls for sensitive AI data and artifacts.
  • Document AI data architectures, schemas, and integration patterns.

Skills

Python
Data engineering
AI/ML workloads
SQL
Collaboration
Security & governance

Education

Bachelor's or Master's degree in CS/DS/ML

Tools

Databricks
Apache Spark
MLflow
Neo4j
LangChain
HuggingFace
Vector databases

Job description

Location: Princeton, NJ & NYC, NY (Hybrid)

Role Overview

The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.


Key Responsibilities


  • Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.

  • Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.

  • Develop MCP servers and enable AI data distribution via MCP.

  • Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.

  • Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.

  • Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.

  • Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.

  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.

  • Develop data quality frameworks specific to AI training datasets and validation data.

  • Collaborate with data scientists to optimize data access patterns and feature store implementations.

  • Implement security and compliance controls for sensitive AI training data and model artifacts.

  • Create comprehensive documentation for AI data architectures, schemas, and integration patterns.


Required Skills and Qualifications


  • Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.

  • 5-7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.

  • Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.

  • Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.

  • Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.

  • Experience with feature engineering, data preprocessing techniques, and ML data pipelines.

  • Knowledge of vector databases, embeddings, and similarity search for AI applications.

  • Proficiency in SQL for structured and unstructured data management.

  • Understanding of data governance, model governance, and AI ethics principles.

  • Strong analytical and problem-solving capabilities with attention to data quality.

  • Excellent collaboration skills for working with data scientists, ML engineers, and architects.


Preferred/Nice-to-Have Skills


  • Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.

  • Knowledge of LangChain, HuggingFace, or other GenAI frameworks.

  • Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.

  • Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.

  • Understanding of AI model explainability and interpretability techniques.

  • Experience with A/B testing frameworks for ML model evaluation.

  • Certification in Databricks, AWS, Azure, or GCP AI/ML services.

  • Publications or contributions to open-source ML projects

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Data Engineer
AI Data Engineer

AIT Global inc. • Princeton (NJ)

On-site
USD 120,000 - 180,000
Senior AI Data Engineer/ Architect
Senior AI Data Engineer/ Architect

Altimetrik • New York (NY)

Hybrid
USD 140,000 - 190,000
Lead AI Engineer
Lead AI Engineer

Anblicks • Richardson (TX)

On-site
USD 180,000 - 240,000
Data AI Engineer
Data AI Engineer

Compunnel, Inc. • Columbus (OH)

On-site
USD 100,000 - 130,000
Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks • Dallas (TX)

On-site
USD 150,000 - 190,000
Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks Inc. • Dallas (TX), Northern (KY)

On-site
USD 150,000 - 230,000
AI Data Engineer
AI Data Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 90,000 - 120,000
AI Data Engineer (US)
AI Data Engineer (US)

AVP VIGILANT TECHNOLOGY PVT LTD • San Francisco (CA)

On-site
USD 115,000 - 195,000
Health, dental, and vision benefits
401(k) retirement benefits
Paid time off and holidays
+2
Lead AI Engineer
Lead AI Engineer

RedStream Technology • Lewisville (TX)

On-site
USD 180,000 - 240,000
Data Engineer, Analytics (Ranking, AI)
Data Engineer, Analytics (Ranking, AI)

Meta • Menlo Park (CA)

On-site
USD 180,000 - 260,000
Bonus
Equity
Benefits