Job Summary
We are looking for a strong Databricks Engineer with 47 years of hands‑on experience in designing and building scalable data engineering solutions using Databricks, Apache Spark, and cloud platforms such as Azure or AWS.
The ideal candidate should have strong expertise in PySpark, SQL, data pipelines, Delta Lake, ETL/ELT, and cloud‑based data platforms, with the ability to work independently on complex data engineering problems.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks and Apache Spark.
- Develop high‑performance ETL/ELT workflows using PySpark and SQL.
- Build and optimize Delta Lake tables and implement reliable data processing frameworks.
- Work with cloud services on Microsoft Azure or AWS to build end‑to‑end data solutions.
- Develop data ingestion pipelines from various sources including databases, APIs, files, and streaming systems.
- Optimize Spark jobs, SQL queries, and Databricks workloads for performance and cost.
- Implement data quality, validation, monitoring, and error‑handling mechanisms.
- Work with large and complex datasets and ensure data accuracy, consistency, and availability.
- Implement incremental data processing and CDC‑based data pipelines where required.
- Collaborate with data architects, analysts, data scientists, and application teams.
- Follow engineering best practices including version control, CI/CD, testing, documentation, and code reviews.
- Troubleshoot production data pipelines and resolve performance and data‑quality issues.
Required Skills
Databricks & Spark
- Strong hands‑on experience with Databricks.
- Strong knowledge of Apache Spark and PySpark.
- Experience with Delta Lake, Delta tables, optimization, and partitioning.
- Understanding of Spark performance tuning, joins, caching, partitioning, and file optimization.
- Experience developing production‑grade notebooks, workflows/jobs, and data pipelines.
Programming & Data
- Strong Python/PySpark programming skills.
- Strong SQL skills, including complex queries, joins, CTEs, window functions, and query optimization.
- Strong understanding of data warehousing and dimensional modeling concepts.
- Experience with batch and, preferably, streaming data processing.
- Understanding of ETL/ELT frameworks and data integration patterns.