Data Engineer

Infinityquest It Services

India

Remote

INR 1,200,000 - 2,100,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Infinityquest It Services is seeking an experienced Data Engineer to modernize our data platform by migrating ETL pipelines to a Databricks-based platform on AWS. You will analyze existing pipelines, design scalable data workflows, and implement PySpark/SQL transformations to ensure data accuracy and performance.

You will collaborate with analytics, product, and engineering teams to understand data requirements, implement CI/CD for data workflows, and maintain observability with monitoring and

Qualifications

  • Strong proficiency in SQL and relational databases (Oracle, PostgreSQL, etc.).
  • Hands-on experience with ETL/ELT pipeline development and migration projects.
  • Strong programming skills in Python and/or Scala.
  • Good experience with Apache Spark (preferably PySpark).
  • Hands-on experience with Databricks (Delta Lake, notebooks, jobs, workflows).
  • Strong knowledge of AWS data engineering services including AWS Glue, Lambda, S3, EventBridge, SQS.
  • Amazon Redshift, Firehose, CloudWatch.
  • Strong understanding of data modeling, data warehousing concepts, and data lake architectures.
  • Experience working with large-scale datasets and distributed processing.
  • Knowledge of pipeline orchestration and workflow management.

Responsibilities

  • Analyze and understand existing ETL pipelines, data sources, and transformation logic across current platforms.
  • Drive end-to-end migration of data pipelines to the new platform using Databricks.
  • Design, build, and maintain scalable data pipelines for batch and near real-time processing in the new platform.
  • Re-engineer legacy ETL/ELT workflows into optimized, scalable, and modular pipelines using Spark/Databricks.
  • Integrate data from multiple heterogeneous sources and ensure seamless data flow across systems.
  • Implement data transformation logic using PySpark/SQL on Databricks and ensure adherence to best practices.
  • Validate migrated pipelines to ensure data accuracy, reconciliation, and consistency with source systems.
  • Optimize pipelines for performance, scalability, and cost efficiency within cloud environments.
  • Work closely with business stakeholders, analysts, and data teams to understand data requirements and use cases.
  • Support reporting, analytics, and downstream applications by providing reliable and high-quality datasets.
  • Monitor production pipelines, troubleshoot failures, and proactively resolve data-related issues.
  • Ensure proper logging, alerting, and observability for all pipelines (using tools like CloudWatch / Databricks monitoring).
  • Follow data governance, security, and compliance standards during data handling and migration.
  • Implement CI/CD pipelines for data workflows and maintain proper documentation of pipelines and processes.

Skills

SQL & DBs
ETL Pipelines
Python/Scala
Spark/PySpark
Spark/Databricks
AWS Data Services
Data Modeling
Distributed Processing
Workflow Orchestration
CI/CD Tools

Tools

Databricks
AWS Glue
Lambda
S3
EventBridge
SQS
Redshift
CloudWatch

Job description

We are looking for an experienced Data Engineer who will play a key role in modernizing our data platform by analyzing existing ETL pipelines and migrating them to the new platform built on Databricks. The ideal candidate should have strong hands-on experience with data engineering best practices, AWS ecosystem, and Databricks-based transformations.

Key Responsibilities
  • Analyze and understand existing ETL pipelines, data sources, and transformation logic across current platforms
  • Drive end-to-end migration of data pipelines to the new platform using Databricks
  • Design, build, and maintain scalable data pipelines for batch and near real-time processing in the new platform
  • Re-engineer legacy ETL/ELT workflows into optimized, scalable, and modular pipelines using Spark/Databricks
  • Integrate data from multiple heterogeneous sources and ensure seamless data flow across systems
  • Implement data transformation logic using PySpark/SQL on Databricks and ensure adherence to best practices
  • Validate migrated pipelines to ensure data accuracy, reconciliation, and consistency with source systems
  • Optimize pipelines for performance, scalability, and cost efficiency within cloud environments
  • Work closely with business stakeholders, analysts, and data teams to understand data requirements and use cases
  • Support reporting, analytics, and downstream applications by providing reliable and high-quality datasets
  • Monitor production pipelines, troubleshoot failures, and proactively resolve data-related issues
  • Ensure proper logging, alerting, and observability for all pipelines (using tools like CloudWatch / Databricks monitoring)
  • Follow data governance, security, and compliance standards during data handling and migration
  • Implement CI/CD pipelines for data workflows and maintain proper documentation of pipelines and processes
Required Skills & Qualifications
  • Strong proficiency in SQL and relational databases (Oracle, PostgreSQL, etc.)
  • Hands-on experience with ETL/ELT pipeline development and migration projects
  • Strong programming skills in Python and/or Scala
  • Good experience with Apache Spark (preferably PySpark)
  • Hands-on experience with Databricks (Delta Lake, notebooks, jobs, workflows)
  • Strong knowledge of AWS data engineering services including:
  • AWS Glue, Lambda, S3, EventBridge, SQS
  • Amazon Redshift, Firehose, CloudWatch
  • Strong understanding of data modeling, data warehousing concepts, and data lake architectures
  • Experience working with large-scale datasets and distributed processing
  • Knowledge of pipeline orchestration and workflow management
Preferred Skills (Good to Have)
  • Experience with data migration to modern data platforms (e.g., Databricks, Lakehouse architecture)
  • Understanding of CDC (Change Data Capture), Delta Lake, and incremental processing patterns
  • Experience with monitoring tools (Grafana, CloudWatch)
  • Familiarity with CI/CD tools (Azure DevOps, Git-based workflows)
Key Expectations from Candidate
  • Ability to quickly understand existing systems and reverse-engineer data pipelines
  • Strong problem-solving skills to modernize and optimize legacy data processes
  • Ownership mindset to drive migration tasks end-to-end with minimal supervision
  • Strong collaboration and communication skills to work across cross-functional teams

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer – Databricks
Senior Data Engineer – Databricks

Aspire, Jordan • India

On-site
INR 1,500,000 - 2,100,000
Data Engineer
Data Engineer

BSL Consulting • Bhopal

On-site
INR 900,000 - 1,300,000
Data Engineer
Data Engineer

Tekskills • Chennai District

On-site
INR 2,000,000 - 4,000,000
Senior Data Engineer
Senior Data Engineer

Tekskills • Pune District, Bengaluru, Delhi

Hybrid
INR 1,500,000 - 2,100,000
Data Engineering Manager
Data Engineering Manager

Crisil • Mumbai

On-site
INR 3,500,000 - 7,500,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India Pvt Ltd • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

SoftEdge • India

On-site
INR 1,800,000 - 2,400,000
Senior Data Engineer (Databricks & AWS Migration) - 573433
Senior Data Engineer (Databricks & AWS Migration) - 573433

NITYO • India

On-site
INR 700,000 - 1,200,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Databricks Engineer
Databricks Engineer

Impronics Technologies • Gurugram District

On-site
INR 4,000,000 - 7,000,000