Job Purpose
As a Data Scientist at SAAL.ai, you will apply machine learning and deep learning techniques to solve real-world business problems across numeric and text-based analytics. You will work across the full data science lifecycle-from data acquisition and feature engineering to model development and deployment-collaborating closely with engineering, product, and DevOps teams to deliver customer-centric AI solutions.
Key Responsibilities
- Develop and deliver high-quality machine learning and deep learning models across a broad range of analytics use cases.
- Understand the practical scope and constraints of AI models within products and apply appropriate modeling techniques to ensure business relevance and technical robustness.
- Acquire data from diverse sources, performing data exploration and preprocessing on structured and unstructured datasets, conducting feature engineering, evaluating algorithms and architectures, and iteratively refining models to improve performance.
- Identify valuable data sources and automate data collection and preparation processes where possible.
- Build, fine-tune, and maintain machine learning pipelines, integrate algorithms and packages into the platform, and contribute new models to the platform marketplace.
- Collaborate closely with engineering and DevOps teams to support model deployment and operationalization.
- Contribute to applied research activities, propose data-driven solutions to business challenges, and work with product, support, and client-facing teams to ensure successful rollout of AI solutions into trials and general availability.
Key Skills
- Strong foundation in the theory and applied practice of machine learning and deep learning. Experience developing pipelines for structured and unstructured data.
- Hands-on experience with deep learning frameworks such as TensorFlow and PyTorch, and machine learning libraries such as scikit-learn.
- Proficiency in Python, with working knowledge of Java or R.
- Experience developing and exposing models through RESTful APIs and containerized deployments using Docker.
- Strong coding practices, including appropriate use of data structures and clean, maintainable code.
- Applied experience in NLP, including the use of transformer-based models and libraries such as Hugging Face, spaCy, or Gensim. Familiarity with SQL, Pandas, Apache Spark, and data processing at scale.
- Experience using ML lifecycle tools such as MLflow.
- Strong analytical thinking, communication skills, and the ability to present technical concepts clearly to internal and external stakeholders.
- Ability to work collaboratively across multidisciplinary teams.
Desirable Skills
Experience working with large language models (LLMs), including offline deployment scenarios. Exposure to cloud-based environments (AWS, GCP, Azure). Experience with data visualization tools and collaborative environments such as Jupyter Notebooks and Git-based workflows. Domain experience with education or finance datasets is an advantage.