Get more replies from employers
Send a job-specific resume in minutes.
Covetus is seeking a Data Engineer to design and optimize ETL/ELT pipelines on Azure Databricks using PySpark. You will build scalable ingestion workflows from diverse sources and implement Delta Lake Bronze-Silver-Gold layers with robust governance.
Responsibilities include developing reusable notebooks, integrating with Azure services, and applying DevOps practices to ensure performance and cost efficiency. Local candidates in NYC/Charlotte preferred, with in-person interviews.
Only local candidates to NYC, NY & Charlotte, NC
Only US Citizens and GC
1 In-Person Interview
Key Responsibilities:
Data Engineering Pipeline Development Design develop and optimize ETLELT pipelines using Azure Databricks PySpark
Build scalable data ingestion workflows from various structured and unstructured sources
Implement transformation logic data cleansing enrichment and validation frameworks
Work with Delta Lake to build medallion architecture Bronze Silver Gold layers
Develop reusable Databricks notebooks and jobs for production data workflows Azure Cloud Integration
Build and orchestrate pipelines using Azure Data Factory ADF
Integrate Databricks with other Azure servicesADLS Azure SQL Event Hub Key Vault Synapse
Optimize compute environments clusters pools autoscaling Implement DevOps processes using Git CICD Azure DevOps Performance Quality Governance
Optimize PySpark jobs for performance and cost efficiency
Implement best practices for data governance security and access control
Troubleshoot production issues and perform rootcause analysis
Conduct code reviews ensuring coding standards and data quality Collaboration Documentation
Work with Data Architects to define architecture and design patterns
Prepare technical documents solution diagrams and runbooks
Collaborate with business stakeholders to understand requirements and translate them into technical solutions
Mandatory Skills:
Azure Databricks notebooks jobs workflows Delta Lake PySpark dataframes Spark SQL optimization debugging Azure Data Factory ADF triggers pipelines integration runtime Data Lake Storage ADLS Gen2 folder structures
Partitioning security CICD Git branching strategies Azure DevOps pipelines SQL
Strong proficiency in writing optimized queries