Sr data engineer

VectorMatch

Hyderabad

On-site

INR 1,800,000 - 2,400,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

VectorMatch in India seeks a Data Engineer to design, build, and maintain robust cloud data pipelines on Azure, using ADF, Databricks, PySpark, and Spark SQL. You will implement the Medallion Architecture across Bronze, Silver, and Gold layers to enable scalable analytics.

You will optimize Spark workloads, ensure data quality, and build CI/CD pipelines for Databricks with Git and Azure DevOps, while solving performance issues like data skew and small-file proliferation.

Qualifications

  • Extensive hands-on experience as a Data Engineer within the Azure ecosystem.
  • Strong programming proficiency in PySpark and Spark SQL.
  • Deep expertise in Delta Lake operations and table management.
  • Advanced SQL skills with complex window functions for analytics.

Responsibilities

  • Design, build, and maintain robust cloud data pipelines on Azure.
  • Implement the Medallion Architecture across Bronze, Silver, and Gold layers.
  • Optimize Spark workloads and manage memory to prevent OOM errors.
  • Build CI/CD pipelines for Databricks using Git, Azure DevOps, and DABs.

Skills

Azure Data Factory
Data Bricks
Data Engineer
SQL
Pyspark
Azure

Tools

Databricks
Azure DevOps
Git
DABs

Job description

Job Description:

Design, build, and maintain robust cloud data pipelines using Azure Data Factory (ADF), Azure Databricks, PySpark, and Spark SQL.

  • Implement and manage the Medallion Architecture — moving and transforming data through Bronze (raw/audit), Silver (cleansing/dedup/SCD), and Gold (business aggregates) layers.
  • Perform complex data transformations, cleansing, deduplication, and incremental loads using Delta MERGE, supporting both Slowly Changing Dimensions (SCD Type 1 and Type 2).
  • Optimize Spark workloads by tuning shuffle partitions, managing memory to prevent OOM errors, and leveraging Adaptive Query Execution (AQE).
  • Apply optimized join strategies, including broadcast joins for small datasets and salting techniques to handle data skew.
  • Implement robust exception handling, file dependency validation, and asynchronous batch processing to ensure pipeline reliability.
  • Ensure high data quality through schema enforcement, schema evolution handling, and validation against expected target criteria.
  • Monitor, troubleshoot, and resolve production job failures by analysing cluster scaling behaviour, Spark UI metrics, and physical query plans.
  • Build and maintain CI/CD pipelines for Databricks using Git, Azure DevOps, and Databricks Asset Bundles (DABs), with environment-specific parameterization for Dev, Test, and Prod.
Mandatory Skills

Azure Data Factory,Data Bricks,Data Engineer,SQL,Pyspark,Azure

Required Qualifications & Skills
  • Extensive hands-on experience as a Data Engineer within the Azure ecosystem.
  • Strong programming proficiency in PySpark and Spark SQL.
  • Deep expertise in Delta Lake operations (ACID transactions, Time Travel, VACUUM, Deep/Shallow Clone) and Managed vs. External table management.
  • Advanced SQL skills, particularly complex window functions (ROW_NUMBER, RANK, DENSE_RANK, LEAD, LAG) for analytics and deduplication.
  • Proven experience with orchestration and monitoring using ADF and Azure Monitor.
  • Solid grounding in software engineering best practices, including version control (Git) and environment-based deployment.
  • Strong analytical and problem-solving skills, with a track record of diagnosing and resolving performance issues such as small-file proliferation, data skew, and excessive driver-side collection.
Preferred & Advanced Skills
  • Experience with Structured Streaming and integrating micro-batches (foreachBatch) into Delta targets.
  • Experience generating deterministic business/composite hashes for records lacking natural primary keys.
  • Familiarity with Infrastructure as Code for deploying ADF and Databricks resources.
  • Exposure to Unity Catalog for data governance, access control, and lineage across workspaces.
  • Familiarity with cost optimization practices — cluster right-sizing, auto-termination policies, and job cluster vs. all-purpose cluster tradeoffs.
  • Experience with data quality frameworks for automated validation.
  • Working knowledge of Python packaging/testing (pytest, unit testing for PySpark transformations).
Location

Chennai / Hyderabad / Bengaluru / Pune / New Delhi

Experience

3 to 12 years

Skills

Azure Data FactoryAzure Data BricksPysparkSpark SQLAzureAzure DevOpsDatabricks

Good to have

Databricks Asset BundlesPython

Requirements:

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Arrow Electronics India Pvt Ltd • Ahmedabad District

On-site
INR 2,800,000 - 4,200,000
Data Engineer(Azure, ADF, Databricks, PySpark, Unity Catalog, SQL) || 6 Years || Bangalore
Data Engineer(Azure, ADF, Databricks, PySpark, Unity Catalog, SQL) || 6 Years || Bangalore

Innova ESI • Bengaluru

On-site
INR 800,000 - 1,600,000
Data Engineer(Azure, ADF, Databricks, PySpark, Unity Catalog, SQL) || 6 to 10 Years || Bangalore
Data Engineer(Azure, ADF, Databricks, PySpark, Unity Catalog, SQL) || 6 to 10 Years || Bangalore

Innova ESI • Bengaluru

On-site
INR 800,000 - 1,500,000
Azure Data Engineer
Azure Data Engineer

Advance Career Solutions • Pune District

On-site
INR 4,000,000 - 6,000,000
Azure Data Engineer | 7+ Years | Bangalore
Azure Data Engineer | 7+ Years | Bangalore

Apexon • Bengaluru

On-site
INR 1,500,000 - 2,800,000
Data Architect
Data Architect

Lancesoft • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Data Engineer
Senior Data Engineer

USEReady • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Azure Data Engineer/ Architect
Azure Data Engineer/ Architect

Capgemini • Gurugram District, Chennai District, Bengaluru

On-site
INR 1,500,000 - 2,700,000
IN_Senior Associate_Azure Data Bricks_Digital Integration_Advisory_Kolkata
IN_Senior Associate_Azure Data Bricks_Digital Integration_Advisory_Kolkata

Price Waterhouse Cooper LLP • Kolkata District

On-site
INR 900,000 - 1,500,000
Azure Databricks Data Engineer
Azure Databricks Data Engineer

Larsen & Toubro Infotech Ltd (LTI) • Bengaluru

On-site
INR 2,500,000 - 4,500,000