AI Data Engineer

EXL

Gurugram District

On-site

INR 1,200,000 - 2,200,000

Full time

10 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

EXL is seeking an AI Data Engineer to design and build robust data pipelines powering AI/ML and GenAI applications in Gurgaon. You will bridge traditional data engineering with modern AI infrastructure to make data AI-ready, scalable, and production-grade.

You will design, implement, and optimize ETL/ELT pipelines using Python, SQL, and PySpark, deploying on Databricks and cloud platforms. This role requires collaboration with Data Scientists and ML Engineers to operationalize models and ensure

Qualifications

  • 3-6 years of experience in Data Engineering.
  • Strong hands-on expertise in SQL and Python.
  • Proficiency in PySpark for large-scale data processing.
  • Working experience with Databricks (Delta Lake, Unity Catalog, notebooks).
  • Exposure to AI/ML data pipelines - vector databases, embeddings, or RAG architecture is a plus.
  • Experience with cloud data platforms (Azure Data Factory, AWS Glue, or equivalent).
  • Understanding of data modeling, warehousing, and pipeline orchestration (Airflow/ADF).
  • Familiarity with LLM ecosystems (LangChain, LlamaIndex) is an added advantage.

Responsibilities

  • Design, build, and optimize ETL/ELT pipelines using Python, PySpark, and SQL.
  • Develop and maintain scalable data pipelines on Databricks (Delta Lake, notebooks, workflows, cluster optimization).
  • Build data infrastructure to support AI/ML and GenAI use cases - including feature engineering, embeddings, and vector data pipelines.
  • Prepare, clean, and structure data (structured & unstructured) for LLM/RAG-based applications.
  • Collaborate with Data Scientists and ML Engineers to operationalize models and AI pipelines.
  • Ensure data quality, governance, and performance across pipelines.
  • Work with cloud platforms (Azure/AWS/GCP) for data storage, compute, and orchestration.
  • Optimize Spark jobs for performance and cost efficiency.

Skills

SQL
Python
PySpark
Databricks
AI pipelines
Azure
AWS
GCP
Airflow
LangChain
LlamaIndex
Delta Lake
Unity Catalog
Pinecone
FAISS
Chroma

Tools

Delta Lake
Unity Catalog
Databricks Notebooks

Job description

We're looking for an AI Data Engineer to design and build robust data pipelines that power AI/ML and GenAI applications. This role sits at the intersection of traditional data engineering and modern AI infrastructure - you'll be responsible for making data AI-ready, scalable, and production-grade.

Key Roles and Responsibilities
  • Design, build, and optimize ETL/ELT pipelines using Python, PySpark, and SQL
  • Develop and maintain scalable data pipelines on Databricks (Delta Lake, notebooks, workflows, cluster optimization)
  • Build data infrastructure to support AI/ML and GenAI use cases - including feature engineering, embeddings, and vector data pipelines
  • Prepare, clean, and structure data (structured & unstructured) for LLM/RAG-based applications
  • Collaborate with Data Scientists and ML Engineers to operationalize models and AI pipelines
  • Ensure data quality, governance, and performance across pipelines
  • Work with cloud platforms (Azure/AWS/GCP) for data storage, compute, and orchestration
  • Optimize Spark jobs for performance and cost efficiency
Required Skills
  • 3-6 years of experience in Data Engineering
  • Strong hands-on expertise in SQL and Python
  • Proficiency in PySpark for large-scale data processing
  • Working experience with Databricks (Delta Lake, Unity Catalog, notebooks)
  • Exposure to AI/ML data pipelines - vector databases, embeddings, or RAG architecture is a plus
  • Experience with cloud data platforms (Azure Data Factory, AWS Glue, or equivalent)
  • Understanding of data modeling, warehousing, and pipeline orchestration (Airflow/ADF)
  • Familiarity with LLM ecosystems (LangChain, LlamaIndex) is an added advantage
Good to Have
  • Experience in BFSI or analytics-heavy domains
  • Exposure to MLOps tools and CI/CD for data pipelines
  • Knowledge of NoSQL/vector databases (Pinecone, FAISS, Chroma)
What We Offer
  • Opportunity to work on cutting-edge AI/data infrastructure projects
  • Collaborative, fast-paced environment
  • Competitive compensation and growth path into ML/AI architecture role
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Bsri Solutions • Chennai District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
AI Data Engineer II (Business Data Analyst II)
AI Data Engineer II (Business Data Analyst II)

UKG • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer - Data & AI Integration Specialist
Data Engineer - Data & AI Integration Specialist

Difinity Digital • Ernakulam

On-site
INR 1,500,000 - 2,500,000
AI Data Engineer (Azure)
AI Data Engineer (Azure)

Careernet • Kolkata District, Pune District, Mumbai

Hybrid
INR 1,200,000 - 2,200,000
Senior Data Engineer (AI Data Platforms & Automation)
Senior Data Engineer (AI Data Platforms & Automation)

E Solutions • Dadri

On-site
INR 1,200,000 - 2,400,000
Data Platform Engineering Lead - AITDS
Data Platform Engineering Lead - AITDS

Cognizant • Hyderabad

On-site
INR 4,000,000 - 6,400,000
Data Engineer
Data Engineer

Xenonstack • Mohali, Chandigarh

On-site
INR 1,200,000 - 2,100,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000
AI Data Engineer
AI Data Engineer

SG Analytics • Mumbai

Hybrid
INR 1,500,000 - 2,700,000