We are seeking a Data Engineer with 3–5 years of experience, based in Chennai/Hyderabad (On-site). The role involves building and optimizing scalable data pipelines, data lakes, and data warehouses while ensuring data quality, governance, and security across analytics and machine learning use cases.
Accountabilities
- Design, develop, and optimize data pipelines, ETL/ELT processes, and data integration workflows.
- Build and maintain data lakes, data warehouses, and real-time streaming pipelines.
- Transform structured and unstructured data into clean, usable datasets for analytics and ML.
- Collaborate with analytics, product, and engineering teams to define and implement data models.
- Ensure data quality, governance, lineage, and security best practices.
- Monitor and optimize pipeline performance and troubleshoot issues in data flow.
- Document data processes and contribute to data engineering standards and frameworks.
- 3–5 years of experience as a Data Engineer.
- Strong programming skills in Python and SQL.
- Hands-on experience with PySpark, Apache Airflow, or dbt.
- Experience with modern cloud data platforms such as:
- AWS (Redshift, Glue, S3)
- GCP (BigQuery, Dataflow)
- Ability to work with both structured and unstructured data.
- Good understanding of data governance, lineage, and security practices.
- Strong troubleshooting and performance optimization skills.
- Familiarity with real-time data streaming tools (Kafka, Kinesis, Pub/Sub).
- Experience supporting machine learning workflows.
- Exposure to containerization/orchestration tools (Docker, Kubernetes).
- Cloud or data engineering certifications.