About the Role
We are looking for an experienced Data Engineer with 4+ years of hands-on experience in designing, developing, and maintaining scalable data pipelines and data platforms.
The ideal candidate will have strong expertise in Python, SQL, PySpark/Spark, ETL/ELT, workflow orchestration, cloud platforms, and data modeling.
You will work closely with Data Scientists, Analysts, Software Engineers, ML Engineers, and business teams to build reliable, scalable, and production-ready data solutions.
Key Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines.
- Build and optimize ETL/ELT workflows using Python, SQL, and PySpark/Spark.
- Develop data ingestion and transformation pipelines for structured and semi-structured data.
- Work with large-scale datasets and optimize distributed data processing for performance and cost.
- Design and implement data models, schemas, data lakes, and data warehouses.
- Develop and manage workflow orchestration using Apache Airflow or similar tools.
- Implement data quality checks, validation, monitoring, and error-handling mechanisms.
- Work with cloud data services and storage platforms such as AWS, Azure, or GCP.
- Collaborate with Data Scientists, Analysts, ML Engineers, and Product teams to deliver production-ready datasets.
- Troubleshoot pipeline failures and improve reliability, scalability, and performance.
- Follow engineering best practices including Git, CI/CD, testing, documentation, and code reviews.
Required Qualifications
- 4+ years of professional experience in Data Engineering or a closely related field.
- Strong programming skills in Python.
- Strong proficiency in SQL, including complex joins, aggregations, CTEs, and window functions.
- Hands-on experience with Apache Spark / PySpark.
- Strong understanding of ETL/ELT pipelines and data engineering concepts.
- Experience with Apache Airflow or another workflow orchestration tool.
- Professional experience with at least one major cloud platform: AWS, Azure, or GCP.
- Strong understanding of data warehousing and data modeling.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or similar.
- Familiarity with Git and CI/CD practices.
- Strong problem-solving, analytical, and communication skills
Preferred Qualifications
- Experience with Kafka or other streaming technologies.
- Experience with Databricks, Snowflake, BigQuery, Redshift, or similar platforms.
- Knowledge of Delta Lake, Apache Iceberg, or modern Lakehouse architectures.
- Experience with Docker and cloud-based DevOps practices.
- Knowledge of data governance, security, and data quality frameworks.
- Experience building real-time/streaming data pipelines.
- Familiarity with dbt or similar transformation frameworks.
Technology Stack
- Programming: Python, SQL
- Data Processing: Apache Spark, PySpark
- Orchestration: Apache Airflow
- Cloud: AWS / Azure / GCP
- Databases: PostgreSQL, MySQL, SQL Server or similar
- Data Platforms: Snowflake, BigQuery, Databricks, Redshift
- Development: Git, CI/CD, Docker
Benefits
- Opportunity to work on large-scale data engineering projects.
- Exposure to modern cloud, data processing, and data platform technologies.
- Work with experienced engineering and data teams.
- Opportunity to build and optimize production-grade data pipelines.
- Competitive compensation based on experience and skills.
- Hybrid working environment in Pune.
What We're Looking For
We are looking for a Data Engineer who can independently own data engineering projects, make sound technical decisions, and build production-grade data pipelines.
The ideal candidate should be comfortable working with large datasets, distributed processing frameworks, cloud platforms, data warehouses, and cross-functional teams.
Candidates should have strong ownership, problem-solving abilities, attention to data quality, and the ability to work in a fast-paced engineering environment.