AI Data Engineer

Smart IT Frame LLC

Princeton (NJ)

On-site

USD 120,000 - 180,000

Full time

Just now
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Smart IT Frame LLC is seeking an AI Data Engineer to design and optimize data pipelines for AI model metadata, training data lineage, and model performance metrics.

You will build Databricks-based infrastructure with Spark, implement MLflow for tracking and model registry, and collaborate with data scientists to advance feature stores, governance, and scalable AI workflows.

Qualifications

  • Bachelor's or Master's in Computer Science, Data Science, ML, or related field.

Responsibilities

  • Design data pipelines for AI model metadata, training data lineage, and model performance tracking.

Skills

Python
ML Frameworks
Databricks/Spark
ML Lifecycle
SQL
Data Governance
Collaboration
Security & Compliance

Education

Bachelor's or Master's in CS/DS/ML or related field

Tools

MLflow
Databricks
Neo4j (Graph DB)

Job description

The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.

Key Responsibilities
  • Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
  • Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
  • Develop MCP servers and enable AI data distribution via MCP.
  • Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
  • Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
  • Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
  • Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
  • Develop data quality frameworks specific to AI training datasets and validation data.
  • Collaborate with data scientists to optimize data access patterns and feature store implementations.
  • Implement security and compliance controls for sensitive AI training data and model artifacts.
  • Create comprehensive documentation for AI data architectures, schemas, and integration patterns.
Required Skills and Qualifications
  • Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
  • 5-7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.
  • Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
  • Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
  • Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
  • Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
  • Knowledge of vector databases, embeddings, and similarity search for AI applications.
  • Proficiency in SQL for structured and unstructured data management.
  • Understanding of data governance, model governance, and AI ethics principles.
  • Strong analytical and problem-solving capabilities with attention to data quality.
  • Excellent collaboration skills for working with data scientists, ML engineers, and architects.
Preferred/Nice-to-Have Skills
  • Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.
  • Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
  • Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
  • Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
  • Understanding of AI model explainability and interpretability techniques.
  • Experience with A/B testing frameworks for ML model evaluation.
  • Certification in Databricks, AWS, Azure, or GCP AI/ML services.
  • Publications or contributions to open-source ML projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Data Engineer - Permanent - Full Time
AI Data Engineer - Permanent - Full Time

Purple Hires Inc • Princeton (NJ)

Hybrid
USD 150,000 - 210,000
AI Data Engineer
AI Data Engineer

AIT Global inc. • Princeton (NJ)

Hybrid
USD 120,000 - 180,000
Senior AI Data Engineer/ Architect
Senior AI Data Engineer/ Architect

Altimetrik • New York (NY)

Hybrid
USD 140,000 - 190,000
Artificial Intelligence Data Engineer
Artificial Intelligence Data Engineer

Green Key Resources • New York (NY)

On-site
USD 150,000 - 190,000
Data AI Engineer
Data AI Engineer

Compunnel, Inc. • Columbus (OH)

On-site
USD 100,000 - 130,000
AI Data Engineer
AI Data Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 90,000 - 120,000
Senior AI Data Engineer
Senior AI Data Engineer

Karsun Solutions • Herndon (VA)

On-site
USD 165,000 - 180,000
Data Engineer - AI
Data Engineer - AI

Compunnel, Inc. • Pennsylvania

On-site
USD 100,000 - 130,000
Lead AI Engineer - Bengaluru, INDIA
Lead AI Engineer - Bengaluru, INDIA

Vytwo • Dallas (TX)

Hybrid
USD 120,000 - 160,000
AI/ML Engineer
AI/ML Engineer

Winaxis LLC • Dallas (TX)

On-site
USD 120,000 - 160,000