We are seeking a highly motivated Data Engineer with experience in Databricks on Microsoft Azure to join a fast-paced team responsible for modernizing enterprise data platforms supporting healthcare research and analytics.
The primary responsibility of this role is to migrate existing Oracle PL/SQL-based ETL processes to Databricks data pipelines. The successful candidate will analyze existing ETL workflows, redesign them using modern cloud technologies, and develop scalable, high-performance data pipelines while ensuring data quality and reliability. The individual will also support existing production ETL processes during the migration.
Key Responsibilities
- Work with business users and stakeholders to gather requirements and translate them into technical solutions.
- Analyze existing Oracle SQL and PL/SQL ETL processes and migrate them to Databricks using Spark SQL and/or PySpark.
- Design, develop, test, deploy, and maintain scalable data pipelines on Microsoft Azure Databricks.
- Create and maintain technical design, data mapping, and system documentation.
- Validate migrated data with business users and subject matter experts to ensure business-level accuracy
- Implement incremental data processing strategies using Databricks and Delta Lake.
- Analyze and convert complex Oracle PL/SQL transformation logic (packages, procedures, functions) into scalable Databricks pipelines.
- Perform data validation and reconciliation between source (Oracle) and target (Databricks) to ensure data accuracy and completeness during migration.
- Provide ongoing support and maintenance for ETL pipelines post-migration in Databricks, including monitoring, troubleshooting, and enhancements.
- Optimize pipeline performance, troubleshoot production issues, perform root cause analysis, and implement solutions.
- Participate in source control, CI/CD, release management, and production support activities.
- Take ownership of assigned projects, manage priorities, and deliver high-quality solutions on schedule.
Required Qualifications
- Bachelor's degree in computer science, Information Systems, Engineering, or related field.
- 4–6 years of experience in ETL, database development, or data engineering.
- 4+ years of advanced SQL development and query optimization.
- 3+ years of experience developing data pipelines using Databricks on Microsoft Azure.
- Experience with Databricks SQL, Spark SQL, PySpark, Delta Lake, and Azure Data Lake Storage (ADLS).
- Experience migrating legacy ETL processes from Oracle or other relational databases to cloud-based data platforms.
- Experience with Git, source control, and CI/CD deployment processes.
- Strong analytical, troubleshooting, and problem‑solving skills.
- Excellent communication skills and the ability to work independently and collaboratively.
Preferred Qualifications
- Experience in the healthcare, biomedical research, or clinical data domain.
- Experience converting Oracle PL/SQL packages, procedures, and functions into Databricks pipelines.
- Experience with Azure Data Factory, Databricks Workflows, or similar orchestration tools.