Insight Advisors Private Limited | Full time
Qualification: B.E./B.Tech, M.E./M.Tech, MCA, orequivalent in Computer Science/IT or a related field.
Work Mode: Hybrid
Primary Skill Set: Databricks, Python, Spark, SQL
Job Description – Data Engineer
Data Pipeline Development & Operations
- Design,build, and operate scalable and reliable data pipelines on the Databricks platform.
- Developend-to-end data workflows from ingestion through transformation toconsumption.
- Implementerror handling, monitoring, and alerting mechanisms.
- Ensurepipeline reliability, performance, and maintainability.
- OptimizeSpark jobs and cluster configurations for improved performance.
- Managecomplex workflows using Databricks Jobs and Workflows.
Legacy Code Modernization
- Refactorlegacy code and ETL pipelines to PySpark.
- Migratetraditional ETL processes to modern ELT patterns on Databricks.
- Identifyopportunities for code and pipeline optimization.
- Ensuredata integrity and backward compatibility during migrations.
- Documentmodernization approaches and migration playbooks.
- Implementdata quality checks and validation frameworks.
- Designand maintain Delta Lake tables with appropriate optimizationstrategies.
- Developreusable data engineering frameworks and libraries.
- Followbest practices for version control, testing and CI/CD.
- Participatein code reviews and troubleshoot production issues.
- Workclosely with data architects, analysts and business stakeholders.
- Collaboratewith Infrastructure, Applications and Cyber teams.
- Sharetechnical knowledge and best practices.
- Mentorjunior Data Engineers on PySpark and Databricks.
Essential Technical Skills
- StrongData Engineering, ETL/ELT and data pipeline design knowledge.
- Hands-on PySpark, including DataFrames, Spark SQL and performanceoptimization.
- Practicalexperience with Databricks Workspace, Cluster Management, Notebooks andJob Orchestration.
- Knowledgeof Databricks Workspace AI Agent capabilities and integration.
- DataModelling experience – Dimensional Modeling, Data Vault or LakehouseArchitecture.
- Strongunderstanding of Delta Lake, including ACID Transactions, SchemaEvolution and optimization.
- Strong Python skills for data processing and automation.
Additional Technical Skills
- Cloudexperience – Azure, AWS or GCP.
- DataGovernance and Security understanding.
- Knowledgeof Structured Streaming.
- DevOpsand CI/CD practices.
- Versioncontrol using Git.
- DataQuality and Testing methodologies.
Professional Experience
- Minimum 5 years relevant Data Engineering experience.
- Atleast 2–3 years hands-on Databricks experience.
- Provenexperience in legacy ETL/code modernization to modern frameworks.
- Experiencebuilding and maintaining production-scale data pipelines.
- Experienceworking with multiple data sources and formats.
Required Certification – Mandatory
At least one of the following:
Preferred Certifications
- DatabricksCertified Associate Developer for Apache Spark.
- AWSCertified Data Analytics.
- GoogleCloud Professional Data Engineer.
- Otherrelevant Data Engineering / Big Data certifications.
Soft Skills
- Strongproblem-solving and analytical skills.
- Excellentcommunication and technical explanation skills.
- Strongcollaboration and stakeholder management.
- Self-motivatedwith attention to detail.
- Adaptableto changing priorities and technologies.
- Client-focusedapproach with commitment to quality delivery.