Responsibilities
- Build, test, and maintain ETL/ELT data pipelines using Azure Databricks and Apache Spark (PySpark).
- Optimize performance and cost‑efficiency of Spark jobs.
- Ensure data quality through validation, monitoring, and alerting mechanisms.
- Understand cluster types, configuration, and use‑cases for serverless computing.
- Implement Unity Catalog for data governance—design and enforce access control policies, manage data lineage, auditing, and metadata governance.
- Enable secure data sharing across teams and external stakeholders.
- Integrate Databricks with Azure Data Lake Storage, Azure Blob Storage, Azure Event Hub, and other cloud data platforms.
- Implement Delta Lake for scalable, ACID‑compliant storage.
- Automate and orchestrate workflows: develop CI/CD pipelines for data workflows using Azure Databricks Workflows or Azure Data Factory.
- Monitor and troubleshoot failures in job execution and cluster performance.
- Collaborate with Data Analysts, Scientists, and Business Teams to translate business needs into scalable data engineering solutions.
- Pull data from a wide variety of APIs using different strategies and methods.
Required Skills & Experience
- Azure Databricks and Apache Spark (PySpark) – strong experience building distributed data pipelines.
- Python – proficiency in writing optimized and maintainable Python code for data engineering.
- Unity Catalog – hands‑on experience implementing data governance, access controls, and lineage tracking.
- SQL – strong knowledge of SQL for data transformations and optimizations.
- Delta Lake – understanding of time travel, schema evolution, and performance tuning.
- Workflow Orchestration – experience with Azure Databricks Jobs or Azure Data Factory.
- CI/CD & Infrastructure as Code (IaC) – familiarity with Databricks CLI, Databricks DABs, and DevOps principles.
- Security & Compliance – knowledge of IAM, role‑based access control (RBAC), and encryption.
Preferred Qualifications
- Experience with MLflow for model tracking & deployment in Databricks.
- Familiarity with streaming technologies such as Kafka, Delta Live Tables, Azure Event Hub, and Azure Event Grid.
- Hands‑on experience with dbt (Data Build Tool) for modular ETL development.
- Certification in Databricks, Azure is a plus.
- Experience with Azure Databricks Lakehouse connectors for Salesforce and SQL Server.
- Experience with Azure Synapse Link for Dynamics, Dataverse.
- Familiarity with other data pipeline strategies, such as Azure Functions, Fabric, ADF, etc.
Soft Skills
- Strong problem‑solving and debugging skills.
- Ability to work independently and in teams.
- Excellent communication and documentation skills.
Equity Statement
Toppan Merrill is an equal opportunity/affirmative action employer. Qualified individuals, including qualified women, minorities, individuals with disabilities, and veterans, are encouraged to apply.