- W2 candidates only, we cannot consider C2C or 1099 for this opportunity.
- Candidates should be eligible to work for any employer in the United States without needing Visa sponsorship
Position Summary
We are seeking a highly skilled Data Engineer with 7+ years of experience designing, developing, and supporting enterprise-scale data platforms. This role will support a strategic initiative to migrate a legacy Netezza environment to Databricks, enabling next-generation analytics and data-driven decision making across customer-facing retail operations.
The successful candidate will play a critical role in modernizing a highly visible and business-critical customer data platform that supports operations. The ideal candidate brings deep expertise in Databricks, PySpark, data engineering best practices, and cloud-based data architecture, along with experience building scalable ingestion, transformation, testing, and monitoring solutions.
Key Responsibilities
- Provide support to the migration of legacy data assets from Netezza to Databricks.
- Design, develop, and optimize scalable data pipelines using PySpark.
- Build and support batch and streaming ingestion frameworks for enterprise data processing.
- Implement data transformation and conversion strategies to support platform modernization initiatives.
- Develop and maintain data solutions utilizing Databricks Lakehouse architecture and Medallion design patterns.
Establish and enforce engineering best practices, including:
- Coding standards
- Automated testing frameworks
- Monitoring and operational support procedures
- Collaborate with business stakeholders, architects, analysts, and engineering teams to deliver high-quality solutions.
- Perform code reviews and mentor junior team members on data engineering best practices.
- Troubleshoot and resolve performance, scalability, and data integrity issues.
- Drive continuous improvement initiatives across the data engineering ecosystem.
- Support integration testing, regression testing, and production deployment activities.
Required Qualifications
- Bachelor's degree in Computer Science, Information Systems, Engineering, or related field.
- 7+ years of Data Engineering experience in enterprise environments.
- Strong hands-on experience with Databricks (required).
- Advanced proficiency in PySpark development.
- Experience developing both batch and streaming data pipelines.
- Strong SQL and data modeling expertise.
- Experience implementing data quality frameworks, validation processes, and reconciliation strategies.
- Experience building and maintaining automated testing and integration suites.
- Experience creating and supporting CI/CD pipelines for data platforms.
- Strong understanding of cloud-based data architectures and modern analytics platforms.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Experience with Netezza migrations or legacy Netezza environments.
- Experience with Delta Lake, Unity Catalog, and Databricks Workflows.
- Experience working with large-scale retail, customer, or omnichannel data platforms.
- Familiarity with Azure cloud services and modern data integration patterns.
- Experience with infrastructure-as-code and DevOps methodologies.
- Knowledge of data governance, lineage, and metadata management solutions.
Top Priority Skills:
Databricks, PySpark, Data Migration, Test Automation, CI/CD, Data Quality, Medallion Architecture, Streaming & Batch Processing