Lead Data Engineer : GEN AI/ LLM/RAG

Trantor Inc.

Bengaluru

On-site

INR 4,200,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Trantor Inc. is seeking a senior Data Engineer / Technical Lead to architect, build, and optimize scalable data platforms supporting batch and streaming workloads.

You will design data models, pipelines, and retrieval-ready datasets for analytics, reporting, ML, and GenAI initiatives, partnering with data scientists and architects. You will mentor teams, establish best practices, and drive CI/CD, data governance, and security across cloud-native lakehouse environments.

Qualifications

  • 10+ years of hands-on data engineering experience
  • Databricks Lakehouse Platform (Delta Lake, Delta Live Tables, Unity Catalog)
  • Medallion architecture (Bronze/Silver/Gold)
  • PySpark and Spark SQL
  • Scalable ETL/ELT pipelines
  • Python and SQL proficiency
  • Data lakes and data warehouses modeling
  • LLMs and GenAI concepts, embeddings, vector databases
  • Cloud platforms (AWS/Azure/GCP)
  • Data security, governance, lineage, IAM and metadata management
  • Agile/Scrum environments and mentoring
  • CI/CD practices

Responsibilities

  • Lead architecture, development, and optimization of scalable data platforms for batch and streaming workloads.
  • Design, implement, and maintain ETL/ELT pipelines using Databricks Lakehouse Platform.
  • Define data models, schemas, and storage strategies for data lakes and warehouses.
  • Develop and curate high-quality datasets, embeddings, and retrieval-ready data stores for LLM-powered apps.
  • Establish engineering standards, code reviews, and CI/CD pipelines.
  • Build end-to-end data workflows using Airflow or similar tools.
  • Lead migrations to modern cloud-native lakehouse architectures.
  • Optimize performance and cost via partitioning, caching, and tuning.
  • Implement data governance, security, lineage, and access controls.
  • Build monitoring, logging, and data quality frameworks; mentor data engineers.
  • Participate in Agile ceremonies and collaborate with stakeholders.

Job description

About Trantor

Trantor is a technology services company focused on outsourced product development and digital re-engineering. Leveraging our CaptiveCoE™ engagement model, we operate as a seamless extension of our client’s teams to provide rapid scalability with predictable budgets. Founded in 2012, Trantor has worked with customers across Tech, FinTech, Media & Cybersecurity industries. We have centers in the US, India, Canada, and Costa Rica. We are consistently rated as the #1 employer in the region with the ability to attract and retain technical talent. Our commitment to excellence and impactful results has translated to long-term relationships and value for our clients and solution partners.

Job Overview

This role requires a solid understanding of Large Language Models (LLMs) and GenAI data ecosystems, enabling the development of high-quality datasets and retrieval-ready pipelines for AI-powered applications. As a technical leader, you will mentor engineering teams, establish best practices, and collaborate closely with architects, data scientists, and business stakeholders to deliver robust data solutions.

Roles & Responsibilities
  • Lead the architecture, development, and optimization of scalable data platforms supporting both batch and streaming workloads.
  • Design, implement, and maintain ETL/ELT pipelines for data ingestion, transformation, and curation using the Databricks Lakehouse Platform and medallion architecture.
  • Define data models, schemas, and storage strategies for data lakes and data warehouses to support analytics, reporting, machine learning, and GenAI initiatives.
  • Develop and curate high-quality datasets, feature engineering pipelines, and retrieval-ready data stores, including embeddings and vector-based data structures, for LLM-powered applications.
  • Establish engineering standards, coding best practices, code review processes, and CI/CD pipelines to ensure maintainable and reliable solutions.
  • Build and automate end-to-end data workflows using orchestration tools such as Apache Airflow or equivalent platforms.
  • Lead migrations from legacy on-premises or cloud-based data warehouses to modern cloud-native and lakehouse architectures.
  • Optimize performance and cost by implementing effective partitioning, caching, compute tuning, and distributed processing strategies.
  • Implement robust data governance, security, lineage, and access control frameworks aligned with organizational compliance requirements.
  • Build monitoring, logging, alerting, and data quality frameworks to ensure reliability and proactive issue resolution.
  • Mentor and guide data engineers while collaborating with architects, data scientists, and business stakeholders to translate business requirements into scalable technical solutions.
  • Participate in and lead Agile ceremonies, including sprint planning, stand-ups, retrospectives, and technical reviews.
Required Skills
  • 10+ years of hands-on experience in data engineering, including leadership of enterprise-scale data platform initiatives.
  • Strong expertise with the Databricks Lakehouse Platform, including Delta Lake, Delta Live Tables, Databricks Workflows, and Unity Catalog.
  • Proven experience implementing the medallion (Bronze/Silver/Gold) architecture for enterprise data platforms.
  • Deep knowledge of distributed data processing using Apache Spark, including PySpark and Spark SQL.
  • Extensive experience building scalable ETL/ELT pipelines for both batch and streaming data processing.
  • Expert proficiency in Python and SQL for data engineering, transformation, validation, and pipeline development.
  • Strong experience designing and managing data lakes and data warehouses using dimensional and lakehouse modeling techniques.
  • Practical understanding of Large Language Models (LLMs) and GenAI concepts, including prompts, embeddings, vector databases, Retrieval-Augmented Generation (RAG), and supporting data pipelines.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP) and its core data services.
  • Demonstrated success leading migrations from legacy platforms such as Hadoop or traditional data warehouses to modern cloud and lakehouse environments.
  • Strong expertise in distributed computing, data partitioning, and performance optimization techniques.
  • Experience implementing data security, governance, lineage, encryption, IAM, and metadata management.
  • Solid understanding of object-oriented programming principles, software design patterns, and CI/CD practices.
  • Experience working within Agile/Scrum environments and mentoring engineering teams.
  • Excellent analytical, problem-solving, stakeholder management, and communication skills.
Preferred Qualifications
  • Experience developing LLM-powered applications or AI data pipelines using frameworks such as LangChain, LlamaIndex, or similar technologies.
  • Hands-on experience with vector databases, including Pinecone, Weaviate, FAISS, or pgvector.
  • Industry certifications in Databricks, AWS, Azure, or Google Cloud Platform.
  • Experience with streaming technologies such as Spark Structured Streaming or Apache Kafka.
  • Familiarity with Infrastructure as Code (Terraform) and DevOps tools such as Git, Jenkins, or Azure DevOps.
  • Exposure to MLOps and LLMOps practices, including model deployment and lifecycle management.
  • Experience working with business intelligence and visualization tools such as Power BI, Tableau, or Amazon QuickSight.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Large Language Model Architect
Large Language Model Architect

Accenture in India • Maharashtra

On-site
INR 4,000,000 - 6,000,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000
Data Engineer (Databricks + AI)
Data Engineer (Databricks + AI)

Durapid Technologies Pvt Ltd • Bengaluru Urban

On-site
INR 1,500,000 - 2,300,000
Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Databricks Data Architect
Databricks Data Architect

Unison Group • Chennai District

On-site
INR 1,800,000 - 3,000,000
Data Engineer - Lead
Data Engineer - Lead

Iris Software • Dadri

On-site
INR 1,500,000 - 2,500,000
Resident Solution Architect
Resident Solution Architect

Celebal Technologies • Dadri

On-site
INR 3,000,000 - 6,000,000
Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Databricks Data Architect
Databricks Data Architect

Unison Group • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Lead Data Engineer
Lead Data Engineer

Celebal Technologies • Dadri, Jaipur, Bengaluru

On-site
INR 3,000,000 - 7,200,000