novio is looking for a Data Engineer based in Mumbai, Maharashtra. The ideal candidate should have over 3 years of hands-on experience, with strong skills in Python, PySpark, and AWS services. Responsibilities include designing ETL pipelines, managing orchestration tasks with Apache Airflow, and developing dashboards using Power BI. The position requires immediate joining and offers a collaborative work environment with other data teams.
Qualifications
3+ years of hands-on experience as a Data Engineer.
Strong proficiency in Python and PySpark programming.
In-depth knowledge of ETL processes and data pipeline architectures.
Responsibilities
Design, build, and maintain scalable ETL/ELT pipelines.
Implement workflow automation using Apache Airflow or Step Functions.
Build and optimize data infrastructure using AWS services.
Integrate and manage data warehouses.
Develop dynamic dashboards to present insights.
Skills
Data Engineering
Python
PySpark
ETL Processes
AWS Services
CI/CD Pipelines
Tools
Apache Airflow
Power BI
Metabase
Superset
Kubernetes
Job description
Job Role
Data Engineer
Location
Mumbai, Andheri
Work Schedule
WFO (Monday-Friday). Looking for immediate joiner within 15 days.
Required Skills
3+ years of hands‑on experience as a Data Engineer
Strong proficiency in Python and PySpark programming for data engineering tasks
In‑depth knowledge of ETL processes and data pipeline architectures
Experience with Airflow or Step Functions for orchestration and scheduling
Solid experience working with AWS services
Proficiency in building and maintaining CI/CD pipelines
Experience with data modelling, database design, and querying
Strong problem‑solving skills, ability to troubleshoot and optimize complex data pipelines
Bonus Skills
Knowledge of cloud platforms (AWS, GCP)
Experience with containerization and Kubernetes
Experience in data security, encryption, and compliance best practices
Strong communication and collaboration skills
Responsibilities
Design, build, and maintain scalable and efficient ETL/ELT pipelines to process large‑scale datasets
Implement and manage workflow automation and orchestration using Apache Airflow or Step Functions
Build and optimize data infrastructure using AWS services
Integrate and manage data warehouses
Design and develop dynamic and interactive dashboards using Power BI, Metabase, and Superset to present insights
Develop and maintain CI/CD pipelines for seamless deployment and continuous integration of data solutions
Monitor and troubleshoot data pipeline issues, ensuring data quality and consistency
Leverage best practices for data governance, security, and compliance in cloud environments
Collaborate with data scientists, analysts, and other stakeholders to understand data needs and deliver reliable data solutions