AI Data Engineer (SP)

Staffingine LLC

New York (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Staffingine LLC is seeking an AI Data Engineer to design robust data pipelines for AI model metadata, training data lineage, and performance tracking. You will build scalable data infrastructure on Databricks with Spark, enabling AI data distribution and feature engineering for model training and inference.

You will collaborate with data scientists to optimize data access, implement model versioning and monitoring, and ensure security and governance for sensitive AI data and artifacts across the

Qualifications

  • BS or MS in CS, DS, ML, or related field.
  • 5–7 years data engineering, at least 2 in AI/ML workloads.
  • Proficient in Python; experience with PyTorch, TensorFlow, scikit-learn.
  • Strong Databricks, Apache Spark, distributed ML workflows.
  • Deep ML lifecycle understanding: training, deployment, monitoring.
  • Knowledge of vector databases, embeddings, similarity search.

Responsibilities

  • Design and implement data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
  • Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
  • Develop MCP servers and enable AI data distribution via MCP.
  • Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
  • Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
  • Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
  • Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
  • Develop data quality frameworks specific to AI training datasets and validation data.
  • Collaborate with data scientists to optimize data access patterns and feature store implementations.
  • Implement security and compliance controls for sensitive AI training data and model artifacts.
  • Create comprehensive documentation for AI data architectures, schemas, and integration patterns.

Skills

Python
Collaborative
Analytical thinking
Data governance
AI ethics

Education

Bachelor's or Master's in CS/DS/ML

Tools

Databricks
Apache Spark
PyTorch
TensorFlow
scikit-learn
SQL

Job description

Job Title: AI Data Engineer (SP)
Job Location:
Princeton, NJ & NYC, NY
Job Type: Full-Time

Job Description:
  • Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
  • Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
  • Develop MCP servers and enable AI data distribution via MCP.
  • Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
  • Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
  • Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
  • Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
  • Develop data quality frameworks specific to AI training datasets and validation data.
  • Collaborate with data scientists to optimize data access patterns and feature store implementations.
  • Implement security and compliance controls for sensitive AI training data and model artifacts.
  • Create comprehensive documentation for AI data architectures, schemas, and integration patterns.
Required Skills and Qualifications
  • Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
  • 5-7 years of hands‑on experience in data engineering, with at least 2 years focused on AI/ML workloads.
  • Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
  • Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
  • Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
  • Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
  • Knowledge of vector databases, embeddings, and similarity search for AI applications.
  • Proficiency in SQL for structured and unstructured data management.
  • Understanding of data governance, model governance, and AI ethics principles.
  • Strong analytical and problem‑solving capabilities with attention to data quality.
  • Excellent collaboration skills for working with data scientists, ML engineers, and architects.
Preferred/Nice‑to‑Have Skills
  • Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine‑tuning.
  • Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
  • Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
  • Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
  • Understanding of AI model explainability and interpretability techniques.
  • Experience with A/B testing frameworks for ML model evaluation.
  • Certification in Databricks, AWS, Azure, or GCP AI/ML services.
  • Publications or contributions to open-source ML projects
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Data Engineer
AI Data Engineer

Reuben Cooley Inc. • Jersey City (NJ)

Hybrid
USD 150,000 - 210,000
AI Data Engineer
AI Data Engineer

AIT Global inc. • Princeton (NJ)

On-site
USD 120,000 - 180,000
Senior AI Data Engineer/ Architect
Senior AI Data Engineer/ Architect

Altimetrik • New York (NY)

Hybrid
USD 140,000 - 190,000
AI Data Engineer
AI Data Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 90,000 - 120,000
Data AI Engineer
Data AI Engineer

Compunnel, Inc. • Columbus (OH)

On-site
USD 100,000 - 130,000
Data Engineer, Analytics (Ranking, AI)
Data Engineer, Analytics (Ranking, AI)

Meta • Menlo Park (CA)

On-site
USD 180,000 - 260,000
Bonus
Equity
Benefits
Data Scientist
Data Scientist

Agentic Dream • United States

Remote
USD 120,000 - 170,000
AI Platform Engineer - Atlanta, GA - FTE
AI Platform Engineer - Atlanta, GA - FTE

Optomi • Atlanta (GA)

On-site
USD 100,000 - 130,000
Lead AI Engineer
Lead AI Engineer

RedStream Technology • Lewisville (TX)

On-site
USD 180,000 - 240,000
Lead AI Engineer
Lead AI Engineer

Anblicks • Richardson (TX)

On-site
USD 180,000 - 240,000