Lead Data Engineer (Databricks, PySpark & GCP)

Egen

Hyderabad

On-site

INR 5,500,000 - 7,500,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Performance bonus

Job summary

Egen is seeking a Lead Data Engineer in Hyderabad to design, architect, and implement scalable ETL/ELT data pipelines using Python, PySpark, Databricks and a strong GCP/Azure background. You will own end-to-end data solutions, collaborate with data scientists and analysts, and drive data quality and governance initiatives.

The role requires 10+ years of experience in data engineering, hands-on Spark, SQL proficiency, and experience with cloud data platforms.

Qualifications

  • 10+ years of hands-on experience in Python for backend or data engineering projects.
  • Strong understanding and working experience with GCP cloud services (especially Dataflow, BigQuery, Cloud Functions, Cloud Composer, etc.).
  • Working experience with Azure Data Factory (ADF), Azure Databricks, Azure Data Lake Storage Gen2 (ADLS).
  • Solid understanding of data pipeline architecture, data integration, and transformation techniques.

Responsibilities

  • Design, develop, test, and maintain scalable ETL data pipelines using Python, PySpark, Databricks & GCP/Azure.
  • Architect the enterprise solutions with various technologies like GCP, Azure, Databricks, PySpark, and Spark SQL.
  • Work with GCP services such as Dataflow, BigQuery, Cloud Functions, Cloud Composer, GCS, IAM, and Cloud Run.
  • Implement Delta Lake architecture and Medallion Architecture (Bronze/Silver/Gold).
  • Manage version control with GitHub and participate in CI/CD for data projects.
  • Write complex SQL queries for data extraction and validation from relational databases.

Skills

Python
PySpark
Spark SQL
SQL

Education

Bachelor's degree in Computer Science

Tools

Databricks
Google Cloud Platform (GCP)
Azure Data Factory (ADF)
Azure Databricks
ADLS Gen2
BigQuery
Cloud Composer
Airflow

Job description

Job Overview:

We are looking for a skilled and motivated Lead Data Engineer with strong experience in Python programming, PySpark, Databricks and Google Cloud Platform (GCP) to join our data engineering team. The ideal candidate will be responsible for requirements gathering, designing, architecting the solution, developing, and maintaining robust and scalable ETL (Extract, Transform, Load) & ELT data pipelines. The role involves working with customers directly, gathering requirements, discovery phase, designing, architecting the solution, using various GCP services, implementing data transformations, data ingestion, data quality, and consistency across systems, and post post-delivery support.

Experience Level:

10 to 16 years of relevant IT experience

Key Responsibilities:
  • Design, develop, test, and maintain scalable ETL data pipelines using Python, PySpark, Databricks & GCP / Azure.
  • Architect the enterprise solutions with various technologies like GCP, Azure, Databricks, PySpark and Spark SQL.
  • Work extensively on Google Cloud Platform (GCP) services such as:
    • Dataflow for real-time and batch data processing
    • Cloud Functions for lightweight serverless compute
    • BigQuery for data warehousing and analytics
    • Cloud Composer for orchestration of data workflows (on Apache Airflow)
    • Google Cloud Storage (GCS) for managing data at scale
    • IAM for access control and security
    • Cloud Run for containerized applications
Should have experience in the following areas :
  • Develop production-grade Databricks notebooks and workflows.
  • Build data transformation pipelines using PySpark and Spark SQL.
  • Implement Delta Lake architecture.
  • Design Bronze, Silver, and Gold data layers using the Medallion Architecture.
  • Implement Databricks Workflows/Jobs and dependency management.
  • Tune Spark jobs for large-scale data processing.
  • Optimize cluster configuration and compute utilization.
  • Implement appropriate partitioning, caching, and file-size optimization strategies.
  • Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery.
  • Implement and enforce data quality checks, validation rules, and monitoring.
  • Collaborate with data scientists, analysts, and other engineering teams to understand data needs and deliver efficient data solutions.
  • Manage version control using GitHub and participate in CI/CD pipeline deployments for data projects.
  • Write complex SQL queries for data extraction and validation from relational databases such as SQL Server, Oracle, or PostgreSQL.
  • Document pipeline designs, data flow diagrams, and operational support procedures.
Required Skills:
  • 10+ years of hands-on experience in Python for backend or data engineering projects.
  • Strong understanding and working experience with GCP cloud services (especially Dataflow, BigQuery, Cloud Functions, Cloud Composer, etc.).
  • Working experience with Azure Data Factory (ADF), Azure Databricks, Azure Data Lake Storage Gen2 (ADLS).
  • Solid understanding of data pipeline architecture, data integration, and transformation techniques.
  • Experience in working with version control systems like GitHub and knowledge of CI/CD practices.
  • Experience in Apache Spark, Kafka, Redis, Fast APIs, Airflow, GCP Composer DAGs.
  • Strong experience in SQL with at least one enterprise database (SQL Server, Oracle, PostgreSQL, etc.).
  • Experience with PySpark is required.
  • Experience in data migrations from on-premise data sources to Cloud platforms.
Good to Have (Optional Skills):
  • Experience with AWS services.
Additional Details:
  • Excellent problem-solving and analytical skills.
  • Strong communication skills and ability to collaborate in a team environment.
Education:

Bachelor's degree in Computer Science, a related field, or equivalent experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Coforge • Dadri

On-site
INR 1,800,000 - 2,600,000
Gcp Data Engineer
Gcp Data Engineer

Lloyds Technology Centre • Hyderabad

Hybrid
INR 2,500,000 - 4,200,000
Lead Data Engineer - R01553055
Lead Data Engineer - R01553055

Brillio • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Competitive compensation and benefits
Remote flexibility
Growth opportunities
Lead Data engineer
Lead Data engineer

HMG AMERICA LLC • Bengaluru Urban

On-site
INR 4,000,000 - 7,000,000
Lead Data Engineer
Lead Data Engineer

Experis • Pune District

On-site
INR 2,500,000 - 5,000,000
Gcp Data Engineer
Gcp Data Engineer

Coforge • Greater Noida

On-site
INR 1,500,000 - 2,100,000
Lead Data Engineer
Lead Data Engineer

Cynosure Corporate Solutions • Chennai District

On-site
INR 400,000 - 600,000
Opportunity for Data Engineer
Opportunity for Data Engineer

Hinduja Tech Limited • Pune District

On-site
INR 1,800,000 - 3,000,000
ETL Technical Product Owner / Manager
ETL Technical Product Owner / Manager

enGen Global • Chennai District

On-site
INR 1,500,000 - 2,100,000
Hiring For GCP Data Engineer For PAN India
Hiring For GCP Data Engineer For PAN India

HCLTech • Dadri, Chennai District, Bengaluru

On-site
INR 2,500,000 - 4,200,000