A data processing company in Pune is looking for a skilled individual to develop and maintain data pipelines. The role involves collaborating with cross-functional teams to design efficient solutions using Python and relevant libraries. Candidates will focus on automating tasks and improving pipeline efficiency while ensuring code quality through testing and documentation. Experience with scheduling tools like Airflow is essential for this position.
Responsibilities
Develop and maintain data pipelines for processing large datasets using Python.
Parse log files and other data sources to extract relevant insights.
Connect to databases and APIs to fetch and manipulate data.
Design efficient data structures and algorithms for data processing.
Lead projects end-to-end, collaborating with cross-functional teams.
Write clean, maintainable, and scalable code with a focus on testing.
Set up and manage scheduling tools like Airflow for automation.
Continuously improve data pipeline efficiency with best practices.
Job description
Responsibilities
Develop and maintain data pipelines for processing large datasets using Python and relevant libraries such as Pandas and NumPy.
Parse log files and other data sources to extract relevant information and insights.
Connect to various databases and APIs to fetch, manipulate and store data as required.
Design and develop efficient data structures and algorithms for data processing, cleaning, and analysis.
Lead projects end‑to‑end and collaborate with cross‑functional teams, including data scientists, analysts, and product managers.
Write clean, maintainable, and scalable code, with a strong emphasis on testing and documentation.
Set up and manage scheduling tools, such as Airflow, to automate data processing tasks.
Continuously improve the efficiency and effectiveness of data pipelines by identifying and implementing best practices and new technologies.