AI Data Engineer

EXL

Gurugram District

On-site

INR 1,500,000 - 2,800,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EXL in Gurugram, India seeks an AI Data Engineer to design and build robust data pipelines powering AI/ML and GenAI applications, ensuring data is AI-ready, scalable, and production-grade.

You will design, develop, and optimize ETL/ELT pipelines with Python, PySpark, and SQL, and work with Databricks Delta Lake and Unity Catalog to enable AI workflows. Collaboration with data scientists and ML engineers is essential.

Qualifications

  • 3–6 years of experience in Data Engineering.
  • Strong hands-on expertise in SQL and Python.
  • Proficiency in PySpark for large-scale data processing.
  • Working experience with Databricks (Delta Lake, Unity Catalog, notebooks).
  • Exposure to AI/ML data pipelines — vector databases, embeddings, or RAG architecture is a plus.
  • Experience with cloud data platforms (Azure Data Factory, AWS Glue, or equivalent).
  • Understanding of data modeling, warehousing, and pipeline orchestration (Airflow/ADF).
  • Familiarity with LLM ecosystems (LangChain, LlamaIndex) is an added advantage.

Responsibilities

  • Design, build, and optimize ETL/ELT pipelines using Python, PySpark, and SQL.
  • Develop and maintain scalable data pipelines on Databricks (Delta Lake, notebooks, workflows, cluster optimization).
  • Build data infrastructure to support AI/ML and GenAI use cases — including feature engineering, embeddings, and vector data pipelines.
  • Prepare, clean, and structure data for LLM/RAG-based applications.
  • Collaborate with Data Scientists and ML Engineers to operationalize models and AI pipelines.
  • Ensure data quality, governance, and performance across pipelines.
  • Work with cloud platforms (Azure/AWS/GCP) for data storage, compute, and orchestration.
  • Optimize Spark jobs for performance and cost efficiency.

Skills

SQL
Python
PySpark
Databricks
LLM / GenAI pipelines
Cloud platforms (Azure/AWS/GCP)
Airflow/ADF
Data modeling & warehousing

Tools

Delta Lake
Unity Catalog
Airflow
ADF

Job description

We're looking for an AI Data Engineer to design and build robust data pipelines that power AI/ML and GenAI applications. This role sits at the intersection of traditional data engineering and modern AI infrastructure — you'll be responsible for making data AI-ready, scalable, and production-grade.

Key Roles and Responsibilities
  • Design, build, and optimize ETL/ELT pipelines using Python, PySpark, and SQL
  • Develop and maintain scalable data pipelines on Databricks (Delta Lake, notebooks, workflows, cluster optimization)
  • Build data infrastructure to support AI/ML and GenAI use cases — including feature engineering, embeddings, and vector data pipelines
  • Prepare, clean, and structure data (structured & unstructured) for LLM/RAG-based applications
  • Collaborate with Data Scientists and ML Engineers to operationalize models and AI pipelines
  • Ensure data quality, governance, and performance across pipelines
  • Work with cloud platforms (Azure/AWS/GCP) for data storage, compute, and orchestration
  • Optimize Spark jobs for performance and cost efficiency
Required Skills
  • 3–6 years of experience in Data Engineering
  • Strong hands-on expertise in SQL and Python
  • Proficiency in PySpark for large-scale data processing
  • Working experience with Databricks (Delta Lake, Unity Catalog, notebooks)
  • Exposure to AI/ML data pipelines — vector databases, embeddings, or RAG architecture is a plus
  • Experience with cloud data platforms (Azure Data Factory, AWS Glue, or equivalent)
  • Understanding of data modeling, warehousing, and pipeline orchestration (Airflow/ADF)
  • Familiarity with LLM ecosystems (LangChain, LlamaIndex) is an added advantage
Good to Have
  • Experience in BFSI or analytics-heavy domains
  • Exposure to MLOps tools and CI/CD for data pipelines
  • Knowledge of NoSQL/vector databases (Pinecone, FAISS, Chroma)
What We Offer
  • Opportunity to work on cutting-edge AI/data infrastructure projects
  • Collaborative, fast-paced environment
  • Competitive compensation and growth path into ML/AI architecture role
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
AI Data Engineer -Azure
AI Data Engineer -Azure

The Glove • Bengaluru

On-site
INR 1,500,000 - 3,500,000
Data Engineer (Databricks + AI)
Data Engineer (Databricks + AI)

Durapid Technologies Pvt Ltd • Bengaluru Urban

On-site
INR 1,500,000 - 2,300,000
Data and AI Engineer
Data and AI Engineer

Western Digital • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Data Engineering (Immediate)
Data Engineering (Immediate)

UsefulBI Corporation • Bengaluru

On-site
INR 900,000 - 1,800,000
AI-ML Data Engineer
AI-ML Data Engineer

CoreFlex Solutions Inc. • Pune District

On-site
INR 1,000,000 - 1,500,000
AI-ML Data Engineer Job ID: 324609
AI-ML Data Engineer Job ID: 324609

CoreFlex Solutions Inc. • Pune District

On-site
INR 800,000 - 1,200,000
Senior Data Engineer (AI/ML)
Senior Data Engineer (AI/ML)

Neolatika • Maharashtra

On-site
INR 2,000,000 - 3,600,000
Data Scientist
Data Scientist

Cognizant • Bengaluru Urban

On-site
INR 4,000,000 - 7,000,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000