Relocation: Candidates willing to relocate from EST or CST are acceptable
Job Summary
We are looking for a Senior AI Data Engineer with 10+ years of experience in data engineering and modern AI/ML technologies. The ideal candidate will have strong expertise in building scalable data platforms and data lakes, along with hands-on experience supporting Generative AI, RAG, LLM, and MCP-based solutions.
Key Responsibilities
- Design, develop, and maintain scalable data engineering and data lake solutions.
- Build and optimize high-volume data pipelines using Python, Spark, and SQL.
- Prepare and transform structured and unstructured data for AI/ML and Generative AI applications.
- Develop data pipelines and architectures supporting RAG (Retrieval-Augmented Generation) and LLM applications.
- Work with AI/ML teams to integrate data platforms with modern AI solutions.
- Develop and maintain data ingestion, transformation, processing, and orchestration workflows.
- Design data models and data architectures optimized for analytics and AI use cases.
- Implement solutions for data quality, governance, security, scalability, and performance.
- Work with LLMs, embeddings, vector-based retrieval, and RAG architectures.
- Contribute to emerging MCP (Model Context Protocol)-based AI integrations and applications.
- Troubleshoot and optimize data pipelines and distributed processing environments.
- Collaborate with data scientists, AI/ML engineers, architects, and business stakeholders.
Required Skills
- 10+ years of experience in Data Engineering.
- Strong hands-on experience with Python.
- Strong experience with Apache Spark / PySpark.
- Advanced SQL skills.
- Experience designing and working with Data Lakes and modern data platforms.
- Strong understanding of AI/ML data pipelines and architectures.
- Hands-on experience with RAG and LLM-based applications.
- Understanding of MCP (Model Context Protocol) and its application in AI systems.
- Experience working with structured and unstructured data.
- Strong understanding of data modeling, ETL/ELT, data processing, and distributed computing.
- Excellent problem-solving and communication skills.
Preferred Experience
- Experience with Generative AI and LLM ecosystems.
- Experience with vector databases and embedding pipelines.
- Experience integrating enterprise data platforms with AI applications.
- Experience with cloud-based data engineering platforms.
- Experience working in large-scale enterprise environments.
Work Location
This is a hybrid position based in Princeton, NJ or NYC, NY. Candidates must be able to work in the required hybrid model.
Note: Visa-independent candidates only. Relocation candidates from the EST or CST time zones may be considered.