Role- Data Engineer
Location- New York
Job Description
We are seeking an experienced Data Engineer to join our team and build robust, scalable data pipelines. In this role, you will:
- Design and implement scalable PySpark data pipelines for batch and streaming workloads
- Optimize Spark jobs and queries for performance and cost efficiency
- Build and maintain ETL/ELT processes following data engineering best practices
- Troubleshoot and resolve complex data pipeline and processing issues
- Collaborate with data teams to ensure data quality and reliability
Top Skills
Databricks Platform Experience
- Hands-on development experience with Databricks notebooks and workflows
- Proficiency in Python and PySpark for data transformation and processing
- Working knowledge of Unity Catalog for data discovery and lineage
- Experience with cluster configuration and job scheduling
- Delta Lake development and optimization techniques
- Databricks SQL for data analysis and reporting
Data Engineering & Pipeline Development
- Advanced ETL/ELT pipeline design and development
- Delta Lake performance tuning (Z-ordering, data skipping, compaction, vacuuming)
- Real-time streaming data pipelines using Structured Streaming and Delta Live Tables
- Query performance optimization and debugging slow-running jobs
- Data quality validation and testing frameworks
- Incremental data processing patterns (CDC, SCD Type 2)
Data Processing & Optimization
- Spark optimization techniques (partitioning, bucketing, caching, broadcast joins)
- Working with large-scale datasets (terabytes to petabytes)
- Data pipeline orchestration and scheduling
- Monitoring and alerting data pipelines
- Implementing Bronze/Silver/Gold (Medallion) data layer patterns
Required Technical Skills
- Databricks & Spark Proficiency: 3+ years of hands-on experience building data pipelines in Databricks; deep understanding of Spark fundamentals, transformations, actions, and performance optimization techniques including partitioning, caching, and resource management
- Advanced PySpark and SQL Skills: Expert-level proficiency writing production-quality PySpark code and complex SQL queries for data transformation, aggregation, and analysis; experience with DataFrame API, Spark SQL, and UDFs; strong understanding of lazy evaluation and execution plans
- Data Engineering & ETL/ELT: Proven experience building and maintaining production data pipelines; hands-on experience with incremental data loading, change data capture (CDC), and slowly changing dimensions; experience handling data quality issues and implementing data validation frameworks
- Cloud & Big Data Technologies: Strong proficiency with AWS services (S3, EC2, IAM, Glue, Athena); experience working with large-scale distributed data processing; familiarity with data formats (Parquet, Delta, JSON, Avro) and compression techniques
- DevOps & CI/CD: Experience with version control (Git) and CI/CD pipelines using GitLab, GitHub Actions, or similar tools; familiarity with testing data pipelines and deployment automation; experience with Databricks Repos and workspace-level integrations
- Data Governance: Understanding of data lineage, cataloging, and metadata management; experience implementing data quality checks and monitoring; knowledge of data privacy and security best practices in cloud environments (nice to have)
Preferred Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field
- Databricks Certified Data Engineer Associate or Professional certification
- Experience with data orchestration tools (Apache Airflow, Databricks Workflows)
- Strong debugging and problem-solving skills
- Excellent communication skills for technical collaboration