Data Scientist

Morningstar

Mumbai City

Hybrid

INR 1,800,000 - 2,800,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Career development

Job summary

Morningstar is seeking a Data Scientist on the AI & ML (Data Collection) team to own AI-powered extraction from PitchBook reports, news, and content. You’ll apply NLP, ML, LLMs, and RAG to improve data quality, coverage, and timeliness of Morningstar data.

You will lead problem definition, data exploration, model development, evaluation, production deployment, monitoring, and continuous improvement. Collaboration with product, engineers, and domain experts is essential.

Qualifications

  • Bachelor’s or Master’s degree in Data Science, CS, Stats, Math, Economics, Engineering, or related field.
  • 2+ years of experience in applied data science, ML, NLP, or information extraction.
  • Experience taking a data science solution from problem definition to production.

Responsibilities

  • End-to-end data science ownership from discovery to production and continuous improvement.
  • Translate business requirements into measurable data science problems and data strategy.
  • Design and optimize NLP, ML, and LLM solutions for document understanding and extraction.
  • Build extraction workflows with parsing, preprocessing, embeddings, RAG, prompt engineering, fine‑Tuning, agentic approaches.
  • Create evaluation datasets and metrics, perform error analysis to guide improvements.
  • Develop robust components and monitor quality in production with engineers.

Skills

Data science
NLP
Python
SQL
Pytorch/TensorFlow
Hugging Face
LangChain
LLMs / RAG
Model evaluation

Education

Bachelor’s or Master’s in Data Science/CS/Stats/Math/Engineering

Tools

Pandas
NumPy
scikit-learn
PyTorch
TensorFlow

Job description

As a Data Scientist on the AI & ML (Data Collection) team, you will own AI-powered solutions that extract structured information from PitchBook’s reports, news, and other content. You will apply data analysis, machine learning, natural language processing (NLP), and generative AI to improve the quality, coverage, and timeliness of PitchBook data.

You will take end-to-end responsibility for data science initiatives, from problem definition, data exploration, and success metrics through model development, evaluation, production deployment, monitoring, and continuous improvement. Your work may include large language models (LLMs), retrieval-augmented generation (RAG), agentic workflows, and other information-extraction techniques.

You will collaborate with Product Managers, Machine Learning Engineers, Software Engineers, and domain experts to deliver scalable solutions. You will remain accountable for model quality and business impact after launching by evaluating performance, investigating regressions, and guiding improvements through data and experimentation. You will also contribute through peer reviews, reproducible work, documentation, and knowledge sharing.
You will join a multidisciplinary team of Data Scientists and Machine Learning Engineers developing AI and ML capabilities for PitchBook’s data collection pipelines. Data Scientists own the analytical and modeling lifecycle and partner with engineering teams to operationalize, scale, and maintain successful solutions.

Primary Job Responsibilities:

  • End-to-End Data Science Ownership: Own extraction and enrichment problems from discovery through production and continuous improvement. Define the problem, select data and methods, establish success criteria, evaluate results, and monitor outcomes
  • Problem Formulation & Data Strategy: Translate business requirements into measurable data science problems. Explore structured and unstructured data, identify quality and source-variability issues, and define training, validation, test, and labeling requirements with domain partners
  • Model Development & Experimentation: Design and optimize NLP, machine learning, and LLM solutions for document understanding and information extraction
  • Extraction Solution Development: Build extraction workflows using document parsing, preprocessing, chunking, feature engineering, embeddings, RAG, prompt engineering, fine-tuning, and agentic approaches
  • Evaluation & Error Analysis: Create representative evaluation datasets and metrics such as precision, recall, F1, field-level accuracy, coverage, confidence, and business impact. Use error analysis to guide model, prompt, data, and workflow improvements
  • Productionization & Model Ownership: Develop robust, testable model components and partner with ML Engineers to integrate solutions into production. Monitor quality, investigate regressions or drift, and prioritize improvements based on customer and business impact
  • Technical Trade-offs & Quality: Evaluate accuracy, coverage, latency, scalability, robustness, and cost. Recommend approaches using empirical evidence, write maintainable code, and document datasets, assumptions, experiments, limitations, and results
  • Collaboration & Innovation: Partner with Product, Data Collection, Engineering, Platform, and domain teams. Evaluate advances in NLP, generative AI, LLMs, and information extraction, and apply methods that deliver measurable value

Skills and Qualifications:

  • Bachelor’s or Master’s degree in Data Science, Computer Science, Statistics, Mathematics, Economics, Engineering, or a related quantitative field
  • 2+ years of experience in applied data science, machine learning, NLP, or information extraction
  • Demonstrated experience taking a data science or machine learning solution from problem definition and experimentation through production launch and ongoing improvement
  • Experience analyzing large, complex structured and unstructured datasets, including exploration, preprocessing, feature engineering, sampling, labeling, and dataset construction
  • Hands‑on experience developing document intelligence or information‑extraction solutions using techniques such as transformers, embeddings, RAG, LLMs, prompt engineering, fine‑tuning, or agentic workflows
  • Strong understanding of experimental design, statistical reasoning, model evaluation, error analysis, and metrics such as precision, recall, F1, field‑level accuracy, confidence, and coverage
  • Proficiency in Python and SQL, with experience using pandas, NumPy, scikit‑learn, and PyTorch or TensorFlow
  • Experience with Hugging Face, LangChain, or comparable NLP and LLM frameworks; ability to write maintainable, testable model and data‑processing code
  • Familiarity with cloud ML environments, version control, automated testing, model monitoring, containers, or data orchestration tools is beneficial
  • Strong communication and collaboration skills, including the ability to explain model behavior, limitations, trade‑offs, and recommendations; experience with financial data, document intelligence, or large‑scale data collection is a plus

Working Conditions

The job conditions for this position are in a standard office setting. Employees in this position use PC and phones on an ongoing basis throughout the day. Limited corporate travel may be required to remote offices or other business meetings and events.

Morningstar’s hybrid work environment gives you the opportunity to collaborate in-person each week as we’ve found that we’re at our best when we’re purposely together on a regular basis. In most of our locations, our hybrid work model is four days in-office each week. A range of other benefits are also available to enhance flexibility as needs change. No matter where you are, you’ll have tools and resources to engage meaningfully with your global colleagues.

I10_MstarIndiaPvtLtd Morningstar India Private Ltd. (Delhi) Legal Entity

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Morningstar Credit Ratings, LLC • Mumbai

Hybrid
INR 2,500,000 - 4,000,000
Hybrid work model
Limited travel
Sr. Data Engineer, Analytics Engineering
Sr. Data Engineer, Analytics Engineering

Morningstar • Mumbai City

Hybrid
INR 2,500,000 - 6,000,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Latitude • Mumbai

Hybrid
INR 4,000,000 - 6,000,000
Hybrid work model
Equal opportunity employer
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Morningstar • Mumbai

Hybrid
INR 3,500,000 - 7,000,000
Hybrid work model
Global collaboration tools
Sr. Data Engineer, Analytics Engineering
Sr. Data Engineer, Analytics Engineering

Morningstar Credit Ratings, LLC • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Senior Data Engineer- Morningstar
Senior Data Engineer- Morningstar

Morningstar • Navi Mumbai

Hybrid
INR 2,500,000 - 4,000,000
Data Scientist
Data Scientist

Viraaj HR Solutions Private Limited • Hyderabad

Hybrid
INR 1,800,000 - 3,200,000
Hybrid work
Bonuses
Fast promotion path
Data Scientist
Data Scientist

Viraaj HR Solutions Private Limited • Khordha

Hybrid
INR 1,400,000 - 2,400,000
Hybrid work
Bonuses
Fast promotion
+1
Data Scientist
Data Scientist

Viraaj HR Solutions Private Limited • Maharashtra

Hybrid
INR 1,500,000 - 3,500,000
Data Scientist
Data Scientist

Viraaj HR Solutions Private Limited • Ernakulam

Hybrid
INR 800,000 - 1,600,000
Hybrid work flexibility
Kaggle sponsorships
Certification reimbursements