Data Platform Lead

Trinity Mobility

Bengaluru

On-site

INR 2,500,000 - 4,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Trinity Mobility in Bengaluru is seeking a seasoned Data Platform Engineer to lead the design, development, and deployment of scalable data platforms enabling analytics, AI/ML, and BI initiatives.

You will architect and optimize end-to-end data pipelines using Apache Airflow, Spark, Kafka, and lakehouse technologies, ensuring performance, reliability, and governance across enterprise data sources and cloud environments.

Qualifications

  • Experience designing scalable enterprise data platforms and architectures.
  • Proficient in Apache Airflow for workflow orchestration and automation.
  • Strong knowledge of lakehouse architectures (Iceberg/Delta/Hudi).
  • Hands-on with distributed processing (Spark) and real-time analytics (Druid).
  • Familiarity with streaming (Kafka) and cloud platforms (Azure/AWS/GCP).

Responsibilities

  • Lead design, development, and implementation of scalable data platform solutions.
  • Architect and develop Airflow-based workflows for scheduling and monitoring.
  • Design and optimize Spark-based batch and real-time data processing.
  • Develop ETL/ELT frameworks for multi-source data ingestion.
  • Build and optimize Data Lake/Warehouse/Lakehouse architectures.
  • Create high-performance analytical models and OLAP with Kylin.
  • Develop real-time analytics platforms with Druid for dashboards.
  • Integrate streaming pipelines using Kafka or similar.
  • Handle data governance, security, and metadata management.
  • Collaborate with Data Scientists, BI, and DevOps teams.

Skills

Data Platform
Mentoring
Agile teamwork
Leadership

Tools

Apache Airflow
Apache Spark
Python
SQL
Shell scripting
ETL/ELT
Data Warehousing
Data Lakes
Lakehouse
Apache Kylin
Apache Druid
Apache Kafka
Apache Iceberg
Delta Lake
Apache Hudi
PostgreSQL
MySQL
MongoDB
Cassandra
Azure
AWS
GCP
Docker
Kubernetes
Jenkins
GitHub Actions

Job description

Role & responsibilities
  • Lead the design, development, and implementation of scalable enterprise Data Platform solutions supporting analytics, AI/ML, and business intelligence initiatives.
  • Architect and develop robust workflow orchestration pipelines using Apache Airflow for scheduling, monitoring, dependency management, and workflow automation.
  • Design, build, and optimize distributed data processing applications using Apache Spark for large-scale batch and real-time data processing.
  • Develop scalable ETL/ELT frameworks for ingesting structured, semi-structured, and unstructured data from multiple enterprise data sources.
  • Design and implement modern Data Lake, Data Warehouse, and Lakehouse architectures using technologies such as Apache Iceberg, Delta Lake, or Apache Hudi.
  • Build high-performance analytical data models and OLAP solutions using Apache Kylin to enable low-latency multidimensional analytics.
  • Develop and optimize real-time analytics platforms using Apache Druid for interactive dashboards, time-series analytics, and high-speed query performance.
  • Integrate streaming data pipelines using Apache Kafka or similar messaging platforms for real-time data ingestion and processing.
  • Design scalable data models that support reporting, business intelligence, advanced analytics, and machine learning workloads.
  • Optimize Spark applications through partitioning, caching, memory tuning, resource allocation, and query optimization to maximize performance.
  • Configure, monitor, and optimize Apache Airflow environments, including DAG development, scheduling strategies, retries, alerting, logging, and failure recovery mechanisms.
  • Develop reusable data engineering frameworks, metadata-driven pipelines, and common libraries to improve development efficiency.
  • Implement data governance, metadata management, lineage, quality validation, and monitoring across the enterprise data platform.
  • Develop APIs and data services to enable secure and efficient data consumption across applications and business platforms.
  • Integrate data from databases, cloud storage, APIs, IoT platforms, enterprise applications, and third-party systems into centralized data platforms.
  • Design and optimize SQL queries, distributed processing workflows, and data storage strategies for maximum scalability and performance.
  • Deploy and manage data platform components using Docker and Kubernetes, ensuring scalability, reliability, and high availability.
  • Work with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform to design, deploy, and manage cloud-native data engineering solutions.
  • Implement CI/CD pipelines using Git, Jenkins, GitHub Actions, or GitLab CI to automate application deployment and infrastructure changes.
  • Collaborate closely with Data Scientists, BI Developers, Product Managers, DevOps Engineers, and Business stakeholders to deliver enterprise data solutions.
  • Lead architecture discussions, establish coding standards, conduct code reviews, and mentor Data Engineers on best engineering practices.
  • Monitor platform performance, troubleshoot production issues, perform root cause analysis, and drive continuous improvements in platform reliability and efficiency.
  • Ensure platform security, access controls, compliance, backup strategies, and disaster recovery planning.
  • Prepare technical documentation, architecture diagrams, deployment guides, and operational runbooks.
  • Evaluate emerging technologies and recommend enhancements to modernize the organization's data platform ecosystem.
  • Drive technical strategy, innovation, and platform roadmap initiatives while ensuring alignment with organizational business objectives.
Preferred candidate profile
  • 58 years of hands-on experience in Data Engineering, Big Data, or Data Platform development.
  • Strong expertise in Apache Airflow for workflow orchestration, scheduling, monitoring, and pipeline automation.
  • Hands-on experience with Apache Spark (Spark Core, Spark SQL, DataFrames, and performance tuning) for large-scale distributed data processing.
  • Proficiency in Python, SQL, and Shell scripting for developing scalable data pipelines and automation.
  • Strong understanding of ETL/ELT, Data Warehousing, Data Lakes, and modern Lakehouse architectures.
  • Experience designing and implementing scalable, high-performance data platforms capable of handling large volumes of structured and unstructured data.
  • Hands-on experience with Apache Kylin for OLAP analytics and multidimensional data modeling is highly preferred.
  • Experience with Apache Druid for real-time analytics, time-series data processing, and low-latency query performance is an added advantage.
  • Knowledge of Apache Kafka or similar streaming technologies for real-time data ingestion and processing.
  • Familiarity with distributed storage technologies such as Apache Iceberg, Delta Lake, or Apache Hudi.
  • Experience working with relational and NoSQL databases such as PostgreSQL, MySQL, MongoDB, Cassandra, or similar databases.
  • Exposure to cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform.
  • Experience with Docker, Kubernetes, Git, and CI/CD tools such as Jenkins or GitHub Actions is preferred.
  • Strong understanding of data governance, data quality, metadata management, and security best practices.
  • Proven ability to optimize Spark workloads, SQL queries, and large-scale data processing pipelines for performance and scalability.
  • Experience in mentoring junior engineers, participating in architecture discussions, conducting code reviews, and contributing to technical decision-making.
  • Strong analytical, troubleshooting, and problem-solving skills with the ability to resolve complex production issues.
  • Excellent communication, stakeholder management, and collaboration skills, with the ability to work effectively in cross-functional Agile teams.
  • Self-motivated, proactive, and passionate about building scalable enterprise data platforms and adopting modern data engineering technologies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineering Manager
Data Engineering Manager

Good co India • India

On-site
INR 2,400,000 - 5,400,000
Cloud Platform Engineer
Cloud Platform Engineer

Northern Trust • Pune District

On-site
INR 3,500,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

GlobalNodes • Gurgaon

On-site
INR 1,500,000 - 2,100,000
Data Architect
Data Architect

New Era India • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Looking For Senior Platform Engineer -Data Platform
Looking For Senior Platform Engineer -Data Platform

Tech Mahindra • Hyderabad, Pune District

On-site
INR 1,500,000 - 2,100,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

Intercontinental Exchange Holdings, Inc. • Hyderabad

On-site
INR 3,000,000 - 5,500,000
Data Engineer
Data Engineer

Bounteous • Chennai District

On-site
INR 2,000,000 - 3,600,000
Lead Data Platform Engineer
Lead Data Platform Engineer

Gen • Chennai District

On-site
INR 3,600,000 - 6,000,000
Solution Architect - Data Engineer
Solution Architect - Data Engineer

KSB Company • Maharashtra

On-site
INR 800,000 - 1,200,000
Senior/Lead Data Engineer
Senior/Lead Data Engineer

Sourcebae • Chennai District

On-site
INR 2,800,000 - 4,800,000