Data Engineer - Data & AI Integration Specialist

Difinity Digital

Ernakulam

On-site

INR 1,500,000 - 2,500,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Difinity Digital is seeking a Data Engineer to own end-to-end data pipelines, ETL/ELT, and AI data infrastructure, working across cloud and open-source technologies. You will translate client requirements into scalable architectures and drive data governance and security.

You will design and implement scalable ETL pipelines, manage data warehouses, RBAC, and enable AI-driven analytics with vector databases, embeddings, and LLM integrations.

Qualifications

  • 5+ years of hands-on experience in data engineering, data architecture, or enterprise data integration.
  • Proficient in Python, SQL, and PySpark.
  • Experience with cloud data services (AWS, GCP, Azure) and open-source data technologies.
  • Ability to create architecture diagrams and present to clients.
  • Experience with data warehouse management, RBAC, and data governance.

Responsibilities

  • Architecture & Client Engagement: Communicate with clients to gather requirements and draft architecture diagrams for data and AI workloads.
  • ETL & Data Pipeline Design: Own, design, build, and maintain scalable ETL/ELT pipelines across cloud-based and open-source environments.
  • AI & Generative AI Data Infrastructure: Ingest and structure data for vector databases, embeddings, LLMs, and AI integrations.
  • Warehouse Management & Governance: Oversee data warehouse administration and enforce governance, quality checks, and validation.
  • Access Control & Security: Implement and manage RBAC across databases, schemas, and AI datasets.
  • Data Ingestion: Extract, transform, and ingest data from relational/non-relational sources, REST APIs, and flat files.
  • Core Development: Write optimized Python, PySpark, and SQL for large-scale data processing.

Skills

Python
SQL
PySpark
Client engagement
Architecture diagrams

Tools

AWS
GCP
Azure
Data warehouses
REST APIs

Job description

Employment Type: Full-Time - Permanent

Experience Level: Mid-Senior Level

About the Role

We are seeking a skilled Data Engineer with expertise in enterprise data integration, open-source technologies, cloud platforms, and modern AI engineering workflows. In this role, you will take end-to-end ownership of data pipeline design, ETL/ELT execution, data warehouse management, data governance, and AI data infrastructure. You will also serve as a technical bridge to clients by capturing business requirements, translating them into technical solutions, creating architecture diagrams, and supporting AI-driven analytics and application integrations.

Key Responsibilities
  • Architecture & Client Engagement: Communicate directly with clients to gather business requirements, analyze technical constraints, and draft comprehensive architecture diagrams for data and AI workloads.
  • ETL & Data Pipeline Design: Own, design, build, and maintain scalable ETL/ELT pipelines across cloud-based and open-source environments.
  • AI & Generative AI Data Infrastructure: Ingest, structure& unstructured and optimize data for vector databases, embedding generation, LLMs, and internal AI application integration (e.g., chatbots, intelligent search).
  • Warehouse Management & Governance: Oversee data warehouse administration and enforce strict data governance, quality checks, and validation frameworks.
  • Access Control & Security: Implement and manage Role-Based Access Control (RBAC) across all databases, schemas, tables, and AI datasets.
  • Data Ingestion: Extract, transform, and ingest data from relational/non-relational databases, REST APIs, web services, and flat files (CSV, JSON, XML, Excel).
  • Core Development: Write optimized, high-performance code using Python, PySpark, and SQL to process large-scale datasets.
Required Qualifications
  • 5+ years of hands-on experience in data engineering, data architecture, or enterprise data integration.
  • Mandatory expertise in Python, SQL, and PySpark.
  • Proven experience working with cloud-based data services (AWS, GCP, Azure, etc.) alongside open-source data technologies.
  • Practical experience building data pipelines that prepare and serve data for AI/ML workloads, vector databases, or embedding generation.
  • Demonstrated ability to build solution architecture diagrams and present technical concepts clearly to clients and stakeholders.
  • Solid experience in data warehouse management, dimensional modeling (star/snowflake schema), and configuring Role-Based Access Control (RBAC).
  • Strong background in end-to-end ETL/ELT workflow design and data governance enforcement.
  • Excellent client-facing communication skills with a focus on active requirement gathering.
  • Experience extracting and integrating data from major enterprise ERP systems (e.g., SAP, Salesforce, Oracle, Microsoft Dynamics).
Preferred Qualifications (Pluses)
  • Experience or strong familiarity with the Microsoft Azure ecosystem (Azure Data Factory, Azure Synapse, ADLS).
  • Understanding of Microsoft Fabric (Lakehouse/Warehouse paradigms).
  • Hands-on experience building or integrating AI tools, chatbots, or LLM applications.
  • Understanding of Power BI integration and key financial reports/business analytics.
  • Familiarity with AI-assisted software development and code-generation tools (e.g., GitHub Copilot, Cursor).
  • Basic knowledge of MLOps practices, model deployment, and serverless workflow orchestration (Azure Functions, Logic Apps).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineering
Data Engineering

CirrusLabs • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Data Engineer
Data Engineer

Ashley Global Capability Center • Chennai

On-site
INR 7,005,000 - 10,508,000
Data Integration Engineer
Data Integration Engineer

Msys Softsol • Hyderabad, Pune District, Bengaluru

On-site
INR 1,400,000 - 2,400,000
Senior Data Engineer (AI Data Platforms & Automation)
Senior Data Engineer (AI Data Platforms & Automation)

E Solutions • Dadri

On-site
INR 1,200,000 - 2,400,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
AI Data Engineer
AI Data Engineer

EXL • Gurugram District

On-site
INR 1,200,000 - 2,200,000
Azure Data Engineer + AI
Azure Data Engineer + AI

Tata Consultancy Services • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Data Engineering(Data & Integration Engineer)
Data Engineering(Data & Integration Engineer)

CirrusLabs • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Azure Data Engineer
Azure Data Engineer

Cirruslabs • Hyderabad, Mangaluru, Bengaluru

Hybrid
INR 2,500,000 - 5,000,000
Data Engineering Lead
Data Engineering Lead

Kumaran Systems • Hyderabad

On-site
INR 2,800,000 - 4,000,000