Lead Data Engineer

Inferyx

Mhalunge

On-site

INR 3,500,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inferyx in India (Maharashtra) seeks a Lead Data Engineer to design and build scalable data pipelines and integrations that support rapidly increasing data volumes and complexity.

You will collaborate with analytics and business teams to understand needs, create pipelines, and improve data models feeding BI and visualization tools. The role requires hands-on experience with Spark, Databricks, SQL, and cloud data platforms.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • 8+ years of experience in data engineering or related roles.
  • Proficiency in Python, Java, or Scala.
  • Strong SQL skills with relational databases (e.g., MySQL, PostgreSQL).
  • Experience with data warehousing concepts and technologies (e.g., Snowflake, Redshift).
  • Familiarity with big data processing frameworks (e.g., Apache Spark, Hadoop).
  • Hands-on experience with ETL tools and data integration platforms.
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Understanding of data modeling principles and data warehousing design patterns.
  • Excellent problem-solving skills and attention to detail.
  • Strong communication and collaboration skills, with the ability to work effectively in a team environment.

Responsibilities

  • Data Pipeline Development: Design, develop, and maintain scalable batch and streaming data pipelines using Apache Spark (PySpark/Scala) and Databricks.
  • Build end-to-end ETL/ELT workflows for ingesting, transforming, and validating data from diverse source systems while ensuring data accuracy, reliability, and performance.
  • Data Modeling & Analytics Enablement: Design and maintain efficient data models, schemas, and curated datasets that support business analytics, reporting, and visualization tools.
  • Optimize data structures for performance, scalability, and cost across lakehouse and data warehouse platforms.
  • Data Integration: Integrate data from multiple internal and external sources, including relational databases, APIs, flat files, and streaming sources.
  • Ensure seamless and reliable data movement across cloud platforms, data lakes, and analytics systems.
  • Performance Optimization: Identify and resolve performance bottlenecks in Spark jobs, Databricks workloads, and data storage layers.
  • Tune Spark configurations, optimize queries, and improve pipeline efficiency to support large-scale data processing.
  • Data Quality & Governance: Implement data quality checks, validation rules, and governance standards to ensure trustworthy data.
  • Monitor data quality metrics and proactively address data issues in collaboration with stakeholders.
  • Collaboration & Stakeholder Engagement: Work closely with data analysts, data scientists, and business teams to understand requirements and deliver data solutions aligned with business objectives.
  • Partner with platform and cloud teams to ensure architectural consistency and best practices.
  • Documentation & Best Practices: Document data pipelines, data models, and technical designs.
  • Follow best practices for software development, version control, CI/CD, and deployment in distributed data environments.
  • Continuous Improvement: Stay current with emerging data engineering technologies, Spark and Databricks enhancements, and cloud data platform innovations.
  • Drive automation and process improvements to increase reliability, scalability, and developer productivity.

Skills

Spark/Databricks
Python/Java/Scala
SQL
ETL/Data Integration
Data Modeling
Data Warehousing
Cloud Platforms (AWS/Azure/GCP)
Data Quality & Governance
BI/Visualization support
Communication & Collaboration

Education

Bachelor's degree in Computer Science or related field

Tools

Databricks
Apache Spark
MySQL/PostgreSQL
Snowflake/Redshift
Hadoop

Job description

About the Company:

We are a global analytics company, with a mission to empower enterprises to build scalable and robust artificial intelligence & machine learning based applications and solutions. We are a team of data engineers and data scientists helping businesses with actionable intelligence and data driven decisions. The Inferyx platform is an end to end data and analytics platform that lets you disrupt and accelerate with data.


Job Description:

We are looking for a Lead Data Engineer who is responsible for the design and development of scalable data pipelines and integrations to support continual increases in data volume and complexity. Work with analytics and business teams to understand their needs, create pipelines, improve data models that feed BI and visualization tools.


Role & responsibilities
  • Data Pipeline Development: Design, develop, and maintain scalable batch and streaming data pipelines using Apache Spark (PySpark/Scala) and Databricks. Build end-to-end ETL/ELT workflows for ingesting, transforming, and validating data from diverse source systems while ensuring data accuracy, reliability, and performance.
  • Data Modeling & Analytics Enablement: Design and maintain efficient data models, schemas, and curated datasets that support business analytics, reporting, and visualization tools. Optimize data structures for performance, scalability, and cost across lakehouse and data warehouse platforms.
  • Data Integration: Integrate data from multiple internal and external sources, including relational databases, APIs, flat files, and streaming sources. Ensure seamless and reliable data movement across cloud platforms, data lakes, and analytics systems.
  • Performance Optimization: Identify and resolve performance bottlenecks in Spark jobs, Databricks workloads, and data storage layers. Tune Spark configurations, optimize queries, and improve pipeline efficiency to support large-scale data processing.
  • Data Quality & Governance: Implement data quality checks, validation rules, and governance standards to ensure trustworthy data. Monitor data quality metrics and proactively address data issues in collaboration with stakeholders.
  • Collaboration & Stakeholder Engagement: Work closely with data analysts, data scientists, and business teams to understand requirements and deliver data solutions aligned with business objectives. Partner with platform and cloud teams to ensure architectural consistency and best practices.
  • Documentation & Best Practices: Document data pipelines, data models, and technical designs. Follow best practices for software development, version control, CI/CD, and deployment in distributed data environments.
  • Continuous Improvement: Stay current with emerging data engineering technologies, Spark and Databricks enhancements, and cloud data platform innovations. Drive automation and process improvements to increase reliability, scalability, and developer productivity.

Preferred candidate profile

  1. Bachelor's degree in Computer Science, Engineering, or related field.
  2. 8+ years of experience in data engineering or related roles.
  3. Proficiency in programming languages such as Python, Java, or Scala.
  4. Strong SQL skills and experience with relational databases (e.g., MySQL, PostgreSQL).
  5. Experience with data warehousing concepts and technologies (e.g., Snowflake, Redshift).
  6. Familiarity with big data processing frameworks (e.g., Apache Spark, Hadoop).
  7. Hands-on experience with ETL tools and data integration platforms.
  8. Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  9. Understanding of data modeling principles and data warehousing design patterns.
  10. Excellent problem-solving skills and attention to detail.
  11. Strong communication and collaboration skills, with the ability to work effectively in a team environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Lead Data Engineer with ML Experience
Lead Data Engineer with ML Experience

Enable Data Incorporated • Hyderabad

On-site
INR 1,600,000 - 2,500,000
Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Lead Data Engineer
Lead Data Engineer

Cynosure Corporate Solutions • Chennai District

On-site
INR 400,000 - 600,000
Big Data Developer
Big Data Developer

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Lead Data Engineer
Lead Data Engineer

SourcingXPress • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Opportunity for Data Engineer
Opportunity for Data Engineer

Hinduja Tech Limited • Pune District

On-site
INR 1,800,000 - 3,000,000
Data Engineering Manager
Data Engineering Manager

Good co India • India

On-site
INR 2,400,000 - 5,400,000
Lead Data Engineer
Lead Data Engineer

Kanerika Inc • Hyderabad

On-site
INR 4,000,000 - 7,000,000