Job DescriptionJoin a team where data engineering meets real-world impact. In this role, youll help design and deliver scalable data solutions using Databricks and PySpark, enabling teams to turn raw data into trusted, analytics-ready assets. Youll collaborate closely with consultants, data engineers, and stakeholders to understand business needs, build reliable pipelines, and support high-quality releases. This is a great opportunity for someone with 23 years of experience who enjoys solving data challenges, improving performance, and learning modern lakehouse practices. If youre motivated by clean engineering, continuous improvement, and working in a collaborative environment where your contributions are visible and valued, this role offers the right mix of ownership, guidance, and growth.
Roles ResponsibilitiesKey Responsibilities:
- Develop and maintain data pipelines and transformations using Databricks and PySpark.
- Implement scalable ETL/ELT workflows to ingest, cleanse, and curate data for downstream analytics and reporting.
- Optimize Spark jobs for performance and cost by tuning partitions, caching, joins, and cluster configurations.
- Build reusable notebooks and modular code to support consistent development and easier maintenance.
- Perform data validation, reconciliation, and quality checks to ensure accuracy and reliability of datasets.
- Collaborate with cross-functional teams to gather requirements, clarify data definitions, and deliver aligned solutions.
- Support deployments and production operations by troubleshooting failures, analyzing logs, and resolving incidents.
- Contribute to documentation, coding standards, and best practices for Databricks-based development.
Minimum Qualifications:
- Bachelors or Masters degree in BTECH, MTECH, MCA, or MSC (or equivalent).
- 23 years of hands-on experience working with Databricks in data engineering or analytics engineering projects.
- Strong experience in PySpark for building transformations and distributed data processing.
- Solid understanding of data pipeline concepts, data modeling basics, and structured/semi-structured data handling.
- Ability to debug and troubleshoot Spark jobs and collaborate effectively within delivery teams.
Technical RequirementETL, PYSPARK, DATABRICKS, Delta Lake, Spark SQL, Data Modeling, Workflow Orchestration, Performance Tuning
- Experience with Spark optimization techniques and practical performance tuning in Databricks environments.
- Familiarity with Delta Lake concepts such as ACID tables, schema evolution, and incremental processing patterns.
- Exposure to orchestrating workflows and managing dependencies for end-to-end pipeline execution.
- Experience working in agile delivery models with strong ownership of tasks, timelines, and quality outcomes.
Strong communication skills to translate requirements into implementable data solutions and clearly document outcomes.
Educational Requirement
MCA,MSc,MTech,Bachelor of Engineering,BTech
Preferred Skills
Technology->Big Data - Data Processing->PySpark,Technology->Data Engineering->Databricks
Service Line
Data Analytics Unit