Key Responsibilities
Data Engineering & Platform Development
- Design, develop, and maintain scalable data products and engineering solutions.
- Build robust batch and real-time data pipelines for enterprise data platforms.
- Develop and optimize data architectures, ensuring scalability, reliability, and performance.
- Support data lakes, data warehouses, lakehouses, and modern analytical ecosystems.
Data Pipeline & Integration
- Develop and maintain data ingestion frameworks from multiple structured and unstructured data sources.
- Build and manage ETL/ELT pipelines for enterprise analytics and reporting.
- Implement automated data processing workflows and orchestration frameworks.
- Ensure high-quality, secure, and governed data movement across platforms.
Data Modeling & Transformation
- Design and implement scalable data models aligned with business requirements.
- Develop transformation frameworks using SQL, Spark, dbt, AWS Glue, and Airflow.
- Implement data modeling concepts including SCD, schema evolution, and dimensional modeling.
- Optimize data processing performance through tuning and troubleshooting.
Cloud & Data Infrastructure
- Design and operate cloud-based data platforms and infrastructure.
- Support data storage, governance, security, and observability requirements.
- Implement scalable solutions across AWS, Azure, or GCP environments.
- Contribute to platform modernization and cloud transformation initiatives.
Streaming & Advanced Data Processing
- Build and support streaming architectures using Kafka, Flink, or equivalent technologies.
- Develop real-time data processing solutions for operational and analytical use cases.
- Optimize streaming performance, monitoring, and reliability.
DevOps & Automation
- Implement CI/CD pipelines and automated deployment frameworks.
- Follow DevOps best practices for code management, testing, and release management.
- Support infrastructure automation and operational excellence initiatives.
Stakeholder Collaboration
- Partner with Data Science, Analytics, Product, and Business teams to gather requirements.
- Translate business needs into scalable technical solutions.
- Provide technical guidance and support for enterprise data initiatives.
Mandatory Skills
Data Engineering
- Data Pipeline Development
- ETL / ELT Frameworks
- Data Warehousing
- Data Lakes & Lakehouses
- Data Modeling
- Data Architecture
Programming & Query Languages
- Python
- SQL (Expert Level)
- Java / Scala (Preferred)
Big Data & Processing Frameworks
- Apache Spark
- Spark Streaming
- Performance Tuning & Debugging
Data Streaming
- Apache Kafka
- Apache Flink
- Real-Time Data Processing
Cloud Platforms
- AWS (Preferred)
- Azure
- GCP
Data Integration & Orchestration
- AWS Glue
- Airflow
- dbt
- Data Ingestion Frameworks
Data Storage & Formats
- Parquet
- Iceberg
- Data Lake Storage Technologies
DevOps & CI/CD
- Git
- CI/CD Pipelines
- Automated Deployment Frameworks
- DevOps Best Practices
Data Governance & Security
- Data Quality Management
- Data Governance
- Data Privacy & Security
- Metadata Management
- Data Observability
Preferred Skills
- Modern Data Platform Architecture
- Lakehouse Implementations
- Cloud-Native Data Engineering
- Enterprise Analytics Platforms
- Advanced Performance Optimization
- Data Product Development
- Infrastructure as Code (IaC)
Education & Experience
- Bachelors or Master’s degree in Computer Science, Engineering, Information Technology, or a related field.
- 8+ years of experience in Data Engineering, Data Platform Development, or related areas.
- Proven experience building enterprise-scale data platforms and cloud-native data solutions.