Roles and Responsibilities
- Building end-to-end data pipelines for ML models, other data driven solutions such that the pipeline is directly usable for deployment/implementation.
- Building and maintaining data pipelines: data cleaning, transformation, roll-up, pre-processing, etc.
- Building/developing data insight solutions for teams such as Credit, Collections, Distribution, Vigilance, HR, etc.
- Building automation solutions using Python, SQL, Docker, etc. as required.
- Database management and automation.
- Working experience on Linux server to install and configure software (related to data science and AI domain), create services, and basic shell scripting.
- REST API development and management using Docker and related cloud technologies.
Technical Skills
Must have
- High Proficiency in Python coding along with good knowledge of SQL (joins, nested queries, etc.)
- Data analysis experience; ability to identify data points and acquisition mechanisms for structured and unstructured data (text/json/xml) for machine learning pipelines.
- Knowledge of Python libraries such as Pandas, SQLAlchemy (or other Python SQL libraries), and optionally matplotlib, numpy, scipy, scikit-learn, nltk.
- Working knowledge of GIT repositories (Github, Gitlab, etc.)
- Experience in developing REST APIs using Django, Flask, FastAPI, etc. (highly appreciated) and deploying APIs on the cloud with Docker (highly appreciated).
- Hands‑on experience on Linux OS/platform to install software, Python packages, create services, build Docker containers, automate, and run code/scripts.
Data management skill sets
- Ability to understand data models and create ETL jobs using Python scripts.
Python scripts
- Automate regular data acquisition, application process, etc. using Python scripts.
Good to have (must be open to learning if doesn’t have already)
- Should be able to work on problems independently or with minimal support.
- Concept of big data and Spark (PySpark) knowledge.
- Cloud experience (AWS, Azure, GCP).