Data Engineer

Swift Staffing

Pune District

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Swift Staffing seeks a skilled Data Engineer to design, build, and operate scalable data pipelines on the Databricks platform in Pune, India. You will develop end-to-end data workflows, implement monitoring, and ensure reliability across ingestion, transformation, and consumption stages.

The role requires 5+ years in data engineering, hands-on Databricks experience, and a track record refactoring legacy code to modern PySpark frameworks.

Qualifications

  • Minimum 5 years of experience in data engineering or related roles.
  • At least 2-3 years hands-on experience with Databricks platform.
  • Proven track record refactoring legacy code to modern frameworks.
  • Experience building and maintaining production data pipelines at scale.
  • Background working across multiple data sources and formats.
  • Experience in agile development environments.
  • Required certifications: Databricks Data Engineer Associate OR Databricks Data Engineer Professional.

Responsibilities

  • Design, build, and operate scalable data pipelines on the Databricks platform.
  • Develop end-to-end data workflows from ingestion through transformation to consumption.
  • Implement robust error handling, monitoring, and alerting mechanisms.
  • Ensure data pipeline reliability, performance, and maintainability.
  • Optimize pipeline performance through efficient Spark job design and cluster configuration.
  • Manage and orchestrate complex data workflows using Databricks Jobs and workflows.
  • Refactor legacy code and data pipelines to PySpark for improved performance.
  • Migrate traditional ETL processes to modern ELT patterns on Databricks.
  • Collaborate with stakeholders to minimize disruption during code transitions.
  • Implement data quality checks and validation frameworks.
  • Design and maintain Delta Lake tables with optimization strategies.
  • Develop reusable code libraries and frameworks for common data engineering tasks.
  • Follow software engineering best practices including version control, testing, and CI/CD.
  • Troubleshoot and resolve data pipeline issues in production environments.

Skills

PySpark
Databricks
Delta Lake
Python
SQL
Cloud platforms
Structured Streaming
DevOps / CI-CD
Git
Data Modeling
Data Quality

Tools

Databricks
Spark
Git
CI/CD

Job description

Role & responsibilities
Data Pipeline Development & Operations
  • Design, build, and operate scalable and reliable data pipelines on the Databricks platform
  • Develop end-to-end data workflows from ingestion through transformation to consumption
  • Implement robust error handling, monitoring, and alerting mechanisms
  • Ensure data pipeline reliability, performance, and maintainability
  • Optimize pipeline performance through efficient Spark job design and cluster configuration
  • Manage and orchestrate complex data workflows using Databricks Jobs and workflows
Legacy Code Modernization
  • Refactor legacy code and data pipelines to PySpark for improved performance and scalability
  • Migrate traditional ETL processes to modern ELT patterns on Databricks
  • Assess existing codebases and identify opportunities for optimization and modernization
  • Ensure backward compatibility and data integrity during migration processes
  • Document refactoring approaches and create migration playbooks
  • Collaborate with stakeholders to minimize disruption during code transitions
Data Engineering Excellence
  • Implement data quality checks and validation frameworks
  • Design and maintain Delta Lake tables with appropriate optimization strategies
  • Develop reusable code libraries and frameworks for common data engineering tasks
  • Follow software engineering best practices including version control, testing, and CI/CD
  • Participate in code reviews and provide constructive feedback to team members
  • Troubleshoot and resolve data pipeline issues in production environments
Collaboration & Knowledge Sharing
  • Work closely with data architects, analysts, and business stakeholders
  • Collaborate with Infrastructure (Infra), Applications (Apps), and Cyber teams
  • Share knowledge and best practices with Team NCS
  • Mentor junior data engineers on PySpark and Databricks technologies
  • Document technical solutions and maintain comprehensive documentation
Essential Technical Skills
  • Data Engineering: Strong foundation in data engineering principles, ETL/ELT processes, and data pipeline design patterns
  • PySpark: Proven hands-on experience developing data pipelines using PySpark, including DataFrames API, Spark SQL, and performance optimization
  • Databricks Platform: Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration
  • Workspace AI Agent: Knowledge of Databricks Workspace AI Agent capabilities and integration
  • Data Modelling: Experience implementing data models including dimensional modelling, data vault, or lakehouse architectures
  • Delta Lake: Understanding of Delta Lake features including ACID transactions, schema evolution, and optimization techniques
  • Python: Strong Python programming skills for data processing and automation
Additional Technical Skills
  • SQL proficiency for data querying and transformation
  • Experience with cloud platforms (Azure, AWS, or GCP)
  • Understanding of data governance and security best practices
  • Knowledge of streaming data processing (Structured Streaming)
  • Familiarity with DevOps practices and CI/CD pipelines
  • Experience with version control systems (Git)
  • Understanding of data quality frameworks and testing methodologies
Professional Experience
  • Minimum 5 years (Relevant Experience) in data engineering or related roles
  • At least 2-3 years of hands-on experience with Databricks platform
  • Proven track record of refactoring legacy code to modern frameworks
  • Experience building and maintaining production data pipelines at scale
  • Background working across multiple data sources and formats
  • Experience in agile development environments
Required Certifications

Required certifications: mandatory to have at least one certification

  • Databricks Certified Data Engineer Associate OR Databricks Certified Data Engineer Professional
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer (APAC Region)
Senior Data Engineer (APAC Region)

ANRGI TECH • Maharashtra

On-site
INR 2,500,000 - 3,800,000
Senior Data Engineer (APAC Region)
Senior Data Engineer (APAC Region)

Anrgi Tech Private Limited • Pune District

On-site
INR 2,000,000 - 2,800,000
Senior Data Engineer (APAC Region)
Senior Data Engineer (APAC Region)

ANRGI TECH Pvt. Ltd. • Pune District

On-site
INR 1,200,000 - 1,800,000
Senior Data Engineer (Databricks & PySpark)
Senior Data Engineer (Databricks & PySpark)

Experis • Pune District

On-site
INR 4,000,000 - 7,000,000
Databricks Data Engineer
Databricks Data Engineer

Infosys • Pune District, Bengaluru, Delhi

On-site
INR 1,200,000 - 1,800,000
Data Engineer (Azure & Databricks)
Data Engineer (Azure & Databricks)

Lufthansa Technik Services India • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Senior/Lead Data Engineer
Senior/Lead Data Engineer

ICICI Lombard • Mumbai

On-site
INR 2,800,000 - 4,000,000
Data Engineer
Data Engineer

Tekskills • Chennai District

On-site
INR 2,000,000 - 4,000,000
Lead Software Engineer
Lead Software Engineer

Impetus • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Lead Data Engineer
Lead Data Engineer

Experis • Pune District

On-site
INR 2,500,000 - 5,000,000