Data Scientist

Morningstar Credit Ratings, LLC

Mumbai

On-site

INR 2,500,000 - 4,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
Limited travel

Job summary

Morningstar India Private Ltd. in Mumbai is seeking a Data Scientist for the AI & ML (Data Collection) team to own AI-powered extraction solutions across PitchBook content.

You will apply NLP, ML, and LLM techniques to improve data quality, coverage, and timeliness, collaborating with product, ML engineers, and domain experts to deploy, monitor, and iterate production systems.

Qualifications

  • Bachelor's or Master's degree in a quantitative field.
  • 2+ years of applied data science, ML, NLP, or information extraction.
  • Experience taking data science solutions from problem to production.
  • Experience handling large structured/unstructured datasets.
  • Hands-on NLP/LLM techniques with transformers, embeddings, RAG, or prompt engineering.

Responsibilities

  • Own extraction problems from discovery to production and improvement.
  • Translate requirements into measurable data science problems and data strategy.
  • Design and optimize NLP, ML, and LLM solutions for document understanding.
  • Build extraction workflows with parsing, embeddings, RAG, and fine-tuning.
  • Create evaluation datasets and metrics; perform error analysis to improve models.

Skills

Python
SQL
Pandas
NumPy
Scikit-learn
PyTorch/TensorFlow
NLP
LLMs
RAG
Prompt engineering
Document intelligence
Experiment design
Communication

Education

Bachelor's or Master's in Data Science / CS

Tools

Hugging Face
LangChain

Job description

As a Data Scientist on the AI & ML (Data Collection) team, you will own AI-powered solutions that extract structured information from PitchBook's reports, news, and other content. You will apply data analysis, machine learning, natural language processing (NLP), and generative AI to improve the quality, coverage, and timeliness of PitchBook data. You will take end-to-end responsibility for data science initiatives, from problem definition, data exploration, and success metrics through model development, evaluation, production deployment, monitoring, and continuous improvement. Your work may include large language models (LLMs), retrieval-augmented generation (RAG), agentic workflows, and other information-extraction techniques. You will collaborate with Product Managers, Machine Learning Engineers, Software Engineers, and domain experts to deliver scalable solutions. You will remain accountable for model quality and business impact after launching by evaluating performance, investigating regressions, and guiding improvements through data and experimentation. You will also contribute through peer reviews, reproducible work, documentation, and knowledge sharing. You will join a multidisciplinary team of Data Scientists and Machine Learning Engineers developing AI and ML capabilities for PitchBook's data collection pipelines. Data Scientists own the analytical and modeling lifecycle and partner with engineering teams to operationalize, scale, and maintain successful solutions.

Primary Job Responsibilities
  • End-to-End Data Science Ownership: Own extraction and enrichment problems from discovery through production and continuous improvement. Define the problem, select data and methods, establish success criteria, evaluate results, and monitor outcomes
  • Problem Formulation & Data Strategy: Translate business requirements into measurable data science problems. Explore structured and unstructured data, identify quality and source-variability issues, and define training, validation, test, and labeling requirements with domain partners
  • Model Development & Experimentation: Design and optimize NLP, machine learning, and LLM solutions for document understanding and information extraction
  • Extraction Solution Development: Build extraction workflows using document parsing, preprocessing, chunking, feature engineering, embeddings, RAG, prompt engineering, fine-tuning, and agentic approaches
  • Evaluation & Error Analysis: Create representative evaluation datasets and metrics such as precision, recall, F1, field-level accuracy, coverage, confidence, and business impact. Use error analysis to guide model, prompt, data, and workflow improvements
  • Productionization & Model Ownership: Develop robust, testable model components and partner with ML Engineers to integrate solutions into production. Monitor quality, investigate regressions or drift, and prioritize improvements based on customer and business impact
  • Technical Trade-offs & Quality: Evaluate accuracy, coverage, latency, scalability, robustness, and cost. Recommend approaches using empirical evidence, write maintainable code, and document datasets, assumptions, experiments, limitations, and results
  • Collaboration & Innovation: Partner with Product, Data Collection, Engineering, Platform, and domain teams. Evaluate advances in NLP, generative AI, LLMs, and information extraction, and apply methods that deliver measurable value
Skills and Qualifications
  • Bachelor's or Master's degree in Data Science, Computer Science, Statistics, Mathematics, Economics, Engineering, or a related quantitative field
  • 2+ years of experience in applied data science, machine learning, NLP, or information extraction
  • Demonstrated experience taking a data science or machine learning solution from problem definition and experimentation through production launch and ongoing improvement
  • Experience analyzing large, complex structured and unstructured datasets, including exploration, preprocessing, feature engineering, sampling, labeling, and dataset construction
  • Hands-on experience developing document intelligence or information-extraction solutions using techniques such as transformers, embeddings, RAG, LLMs, prompt engineering, fine-tuning, or agentic workflows
  • Strong understanding of experimental design, statistical reasoning, model evaluation, error analysis, and metrics such as precision, recall, F1, field-level accuracy, confidence, and coverage
  • Proficiency in Python and SQL, with experience using pandas, NumPy, scikit-learn, and PyTorch or TensorFlow
  • Experience with Hugging Face, LangChain, or comparable NLP and LLM frameworks; ability to write maintainable, testable model and data-processing code
  • Familiarity with cloud ML environments, version control, automated testing, model monitoring, containers, or data orchestration tools is beneficial
  • Strong communication and collaboration skills, including the ability to explain model behavior, limitations, trade-offs, and recommendations; experience with financial data, document intelligence, or large-scale data collection is a plus
Working Conditions

The job conditions for this position are in a standard office setting. Employees in this position use PC and phones on an ongoing basis throughout the day. Limited corporate travel may be required to remote offices or other business meetings and events. Morningstar's hybrid work environment gives you the opportunity to collaborate in-person each week as we've found that we're at our best when we're purposely together on a regular basis. In most of our locations, our hybrid work model is four days in-office each week. A range of other benefits are also available to enhance flexibility as needs change. No matter where you are, you'll have tools and resources to engage meaningfully with your global colleagues.

I10_MstarIndiaPvtLtd Morningstar India Private Ltd. (Delhi) Legal Entity Morningstar is a global independent investment research and financial data company. Here, you'll help uncover what's hidden, simplify what's complex, and create insights that empower investor success. Company overview Morningstar Development Program Morningstar is strongly committed to creating and preserving equal opportunity for all employees and applicants. We make all employment decisions—including recruitment, hiring, compensation, training, promotion, transfer, discipline, termination, and other personnel matters - without regard to race, color, ancestry, religion, sex, national origin, age, disability, protected veteran status, marital status, sexual orientation, genetic information, citizenship, gender identity and expression, parental status, or other legally protected characteristics or conduct.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Morningstar • Mumbai

Hybrid
INR 1,800,000 - 2,600,000
Hybrid work model (4 days in-office)
Engineering Manager, AI & ML (Data Collection)
Engineering Manager, AI & ML (Data Collection)

Morningstar Credit Ratings, LLC • Mumbai

On-site
INR 3,500,000 - 7,500,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Latitude • Mumbai

Hybrid
INR 4,000,000 - 6,000,000
Hybrid work model
Equal opportunity employer
AI Solutions Engineering & Transformation Manager
AI Solutions Engineering & Transformation Manager

Morningstar Credit Ratings, LLC • Mumbai

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work model
Equal opportunity employer
Flexible benefits
Software Development Engineer
Software Development Engineer

Morningstar Credit Ratings, LLC • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Hybrid work model
Collaboration with global teams
Opportunities for growth and learning
AI Engineer
AI Engineer

Morningstar Credit Ratings, LLC • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Sr. Data Engineer, Analytics Engineering
Sr. Data Engineer, Analytics Engineering

Morningstar Credit Ratings, LLC • Mumbai

Hybrid
INR 1,200,000 - 1,800,000
Software Engineer
Software Engineer

Morningstar • Delhi

On-site
INR 1,400,000 - 2,800,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Morningstar Credit Ratings, LLC • Mumbai

On-site
INR 1,500,000 - 2,500,000
Associate Project Manager
Associate Project Manager

Morningstar Credit Ratings, LLC • Mumbai

On-site
INR 1,400,000 - 2,200,000
Hybrid work environment
Global collaboration