Senior Data Scientist

JiBe ERP

Pakistan

On-site

PKR 2,800,000 - 4,200,000

Full time

41 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JiBe ERP, a cloud-based fully integrated ERP for the shipping industry, is seeking an experienced data scientist/ML engineer to design agentic AI systems and OCR pipelines. You will work on orchestrating multi-step reasoning, tool use, and production-grade ML workflows within a collaborative data science team.

Responsibilities include advancing OCR, information extraction, and RAG capabilities, and integrating with Databricks and MongoDB to support model development and data pipelines.

Qualifications

  • 5+ years of hands-on data science or ML engineering experience.
  • Proven ability to develop agentic AI systems with tool use and multi-agent workflows.
  • Deep OCR and document understanding for varied documents.
  • Strong information extraction and document parsing skills.

Responsibilities

  • Design and develop agentic AI systems orchestrating multi-step reasoning and tool use.
  • Develop OCR solutions for complex document types and accurate data extraction.
  • Contribute to RAG pipelines and data prep workflows with Databricks/MongoDB.

Skills

Python programming
FastAPI
OCR
RAG
LLM workflows
Data science
MongoDB
Databricks
Docker
Kubernetes
LangChain
LangGraph

Tools

MongoDB
Databricks
Docker
Kubernetes
LangChain
LlamaIndex
Haystack

Job description

JiBe is a cloud based fully integrated ERP system for the shipping industry. Our goal is to allow shipping companies to improve productivity, efficiency and safety levels, while reducing costs. JiBe ERP enables increased automation and streamlining of processes, creating pre-defined work flows and reducing the usage of email and paper.

Job Responsibilities
Agentic AI System Development
  • Design and develop sophisticated agentic AI systems that orchestrate multi-step reasoning, decision-making, and tool use to solve complex real-world problems
  • OpenRouter, fine-tuned transformer models, and classical ML models — as tools within agentic workflows
  • Continuously evaluate and incorporate emerging agentic frameworks, patterns, and best practices into the team's solutions
Optical Character Recognition (OCR)
  • Develop and improve advanced OCR solutions addressing highly complex and varied document types
  • Work on document and page classification, determining document types, layouts, and structures as a foundation for downstream processing
  • Design and implement sophisticated information extraction pipelines that identify, parse, and structure data from unstructured or semi-structured documents with high accuracy and reliability
Retrieval Augmented Generation (RAG)
  • Contribute to the development of RAG solutions that leverage data extracted through the team's OCR and information extraction pipelines
  • Collaborate closely with the data engineering team on the design of retrieval pipelines, vector stores, and data preparation workflows that underpin RAG systems
  • Work as an integrated member of a strong, established data science team, contributing expertise while aligning with shared architectural and methodological standards
  • Collaborate closely with data engineers to define data requirements, provide feedback on pipeline outputs, and ensure data consumed from Databricks and MongoDB meets the needs of model development
  • Participate in code reviews, knowledge sharing, and the continuous elevation of the team's technical standards
Experience
  • 5+ years of hands-on experience in applied data science or machine learning engineering, with a strong track record of delivering production-grade solutions
  • Demonstrable experience developing agentic AI systems, including tool use, multi-agent orchestration, or LLM-driven workflows
  • Deep expertise in OCR and document understanding, including experience with complex, real-world documents exhibiting high layout and content variability
  • Strong experience with information extraction techniques, including named entity recognition, structured data extraction, and document parsing
Technical Skills
  • Deep expertise in Python, including advanced concepts such as async/await concurrency, decorators, context managers, metaprogramming, type hinting, and design patterns. Proven ability to write clean, maintainable, and performant code following best practices
  • Strong hands-on experience building scalable, production-grade microservices using FastAPI. Proficiency in API design, dependency injection, middleware integration, async endpoint development, OpenAPI/Swagger documentation, and performance optimization.
  • Solid experience with the ML/AI ecosystem: Hugging Face Transformers, scikit-learn, and PyTorch or TensorFlow.
  • Familiarity with ML-Flow, model serving, containerization (Docker), and orchestration (Kubernetes) for ML workloads.
  • Experience training, fine-tuning, and evaluating transformer-based models as well as classical supervised and unsupervised models.
  • Solid understanding of prompt engineering strategies including few-shot learning, chain-of-thought reasoning, system prompts, and prompt templating.
  • Experience optimizing prompts for accuracy, latency, and cost efficiency including LLM evaluation, and best practices for integrating LLMs into production systems.
  • Hands-on experience with LangChain, LangGraph, and other popular LLM orchestration frameworks (e.g., LlamaIndex, Haystack) for building agentic workflows, RAG pipelines, and complex multi-step LLM applications.
  • Familiarity with OpenRouter or equivalent LLM gateway/routing platforms is a serious advantage.
  • Experience with vector databases (e.g., Milvus, Qdrant, Chroma) for semantic search, retrieval-augmented generation (RAG), and efficient similarity search at scale.
  • Experience with Retrieval Augmented Generation (RAG) — including chunking strategies, embedding models, vector search, and retrieval evaluation — is a serious advantage
  • Familiarity with MongoDB or any other NoSQL database is an advantage
  • Experience working with Databricks or similar large-scale data platforms is an advantage
Soft Skills
  • Strong analytical thinking and the ability to break down complex, ambiguous problems into tractable solutions
  • A collaborative, team-first mindset — comfortable integrating into an existing high-performing team and contributing without the need for top-down direction
  • Clear communication skills, with the ability to discuss technical approaches with both technical peers and non-technical stakeholders
  • A proactive attitude toward learning, with genuine curiosity about the rapidly evolving AI landscape
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Scientist: Agentic AI, OCR & RAG Systems
Senior AI Scientist: Agentic AI, OCR & RAG Systems

JiBe ERP • Pakistan

On-site
PKR 2,800,000 - 4,200,000
AI Engineer
AI Engineer

Tkxel • Lahore

On-site
PKR 2,000,000 - 4,000,000
Senior Software Engineer, Dev
Senior Software Engineer, Dev

IBEX Global • Lahore

On-site
PKR 1,200,000 - 1,800,000
AI Engineer
AI Engineer

Jobtailor • Lahore

On-site
PKR 3,000,000 - 4,000,000
Senior Software Engineer - Data Engineering & AI
Senior Software Engineer - Data Engineering & AI

Devsinc, LLC • Islamabad

On-site
PKR 1,674,000 - 3,348,000
AI Developer — Onsite, Sialkot Office
AI Developer — Onsite, Sialkot Office

FabTechSol • Sialkot

On-site
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Islamabad

On-site
PKR 1,800,000 - 3,000,000
AI Integration Engineer
AI Integration Engineer

Fatima Group • Lahore

On-site
PKR 3,000,000 - 6,000,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Karachi Division

On-site
PKR 2,000,000 - 3,200,000
Senior AI Engineer – Agentic AI & Python
Senior AI Engineer – Agentic AI & Python

Convo • Islamabad

On-site
PKR 250,000 - 420,000