Data Scientist/ NLP

Thompsons HR Consulting Pvt Ltd

Bengaluru

On-site

INR 1,200,000 - 2,200,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Thompsons HR Consulting Pvt Ltd seeks a hands‑on Data Scientist NLP to design, develop, evaluate, and deploy production‑grade NLP and ML solutions for complex text‑driven workflows in Bengaluru.

The ideal candidate will own pipelines, embeddings, vector search, and information retrieval tasks, and work with transformer‑based models across classification, NER, and retrieval use cases. Strong Python/SQL, ML fundamentals, and MLOps experience preferred.

Qualifications

  • Strong Python and SQL programming skills with production experience.
  • Hands-on NLP and ML experience across text analytics and classification tasks.
  • Familiarity with embeddings, semantic search, and information retrieval concepts.
  • Experience with Hugging Face, SentenceTransformers, and transformer‑based models.

Responsibilities

  • Design and develop NLP/ML pipelines to transform noisy text into clean representations for modeling and analytics.
  • Build semantic search and retrieval systems using embeddings, vector databases, and ranking techniques.
  • Develop solutions for candidate ranking, out‑of‑vocabulary handling, and semantic matching.
  • Implement and evaluate supervised and hybrid ML approaches including multi‑output and hierarchical classifications.
  • Work with transformer models for text understanding, extraction, and retrieval use cases.

Skills

Python
SQL
NLP
Machine Learning
Transformers
Embeddings
Semantic Search
Information Retrieval
Text Classification
NER
Clustering
Recall@K
MRR
NDCG

Tools

Hugging Face
SentenceTransformers
Transformer‑based models
Tokenization
Vector Search
ANN search

Job description

Data Scientist NLP Job Overview

We are looking for a

We are looking for a hands‑on Data Scientist specializing in Natural Language Processing (NLP) to design, develop, evaluate, and deploy production‑grade NLP and machine learning solutions for complex, text‑driven workflows.

The ideal candidate should have strong expertise in Python, SQL, NLP, Transformers, embeddings, semantic search, information retrieval, and machine learning, with the ability to take solutions from experimentation through production deployment.

Key Responsibilities
  • Design and develop NLP and machine learning pipelines to process noisy, heterogeneous text data and transform it into clean semantic representations for modeling, retrieval, analytics, and downstream applications.
  • Build and optimize semantic search and retrieval systems using embeddings, vector databases, similarity search, and ranking techniques.
  • Develop solutions for candidate ranking, out‑of‑vocabulary handling, semantic matching, and information discovery.
  • Design, implement, and evaluate supervised and hybrid ML approaches, including:
    • Multi‑output classification
    • Hierarchical classification
    • Named Entity Recognition (NER)
    • Entity extraction and parsing
    • Clustering
    • Rule‑based + ML hybrid systems
  • Work with transformer‑based models for text understanding, classification, similarity, extraction, and retrieval use cases.
  • Perform detailed analysis of relationships and decision boundaries across free‑text fields using:
    • Conditional distributions
    • Entropy
    • Mutual information
    • Directional association
    • Embeddings
    • Predictive ablation studies
  • Design experiments and establish appropriate model evaluation metrics and benchmarks.
  • Compare different models and approaches based on accuracy, performance, scalability, latency, and business impact.
  • Fine‑tune and evaluate transformer models using frameworks such as Hugging Face and SentenceTransformers.
  • Deploy, monitor, troubleshoot, and continuously improve production ML/NLP services.
  • Collaborate closely with platform, backend, data engineering, and product teams to integrate ML solutions into production systems.
  • Communicate technical findings, model performance, and trade‑offs effectively to both technical and non‑technical stakeholders.
Required Skills
  • Strong programming experience in Python and SQL.
  • Hands‑on experience developing production‑grade data pipelines and machine learning workflows.
  • Strong understanding of Natural Language Processing (NLP) and text analytics.
  • Practical experience with one or more of the following:
    • Text Classification
    • Semantic Similarity
    • Text Embeddings
    • Information Retrieval
    • Search & Ranking
    • Clustering
    • Named Entity Recognition (NER)
    • Entity Extraction
  • Hands‑on experience with:
    • Hugging Face
    • SentenceTransformers
    • Tokenization
    • Transformer‑based models
    • Model fine‑tuning
    • Model evaluation
  • Strong understanding of embeddings, vector search, and similarity search.
  • Knowledge of cosine similarity and Approximate Nearest Neighbor (ANN) search methods.
  • Familiarity with information retrieval metrics such as:
    • Recall@K
    • MRR (Mean Reciprocal Rank)
    • NDCG (Normalized Discounted Cumulative Gain)
  • Strong analytical and problem‑solving skills with the ability to design experiments, define evaluation metrics, and interpret model results.
  • Ability to evaluate model trade‑offs and clearly communicate technical findings.
Nice to Have
  • Experience in healthcare, medical imaging, document intelligence, enterprise search, recommendation systems, knowledge retrieval, or routing systems.
  • Knowledge of healthcare and enterprise data standards such as:
    • DICOM
    • PACS/RIS
    • HL7
    • FHIR
  • Experience with MLOps and production ML systems.
  • Experience with cloud platforms and API‑based model deployment.
  • Experience with model serving, monitoring, logging, and performance optimization.
  • Knowledge of Responsible AI, data privacy, and secure handling of sensitive text data.
  • Experience working with vector databases/search platforms and large‑scale retrieval systems.
Preferred Candidate Profile

The ideal candidate will have a combination of NLP expertise, machine learning fundamentals, information retrieval knowledge, and production engineering experience. Candidates with experience building semantic search, embedding‑based retrieval, classification, RAG, document intelligence, or enterprise search solutions will be highly preferred.

Core Skills

Python | SQL | NLP | Machine Learning | Hugging Face | SentenceTransformers | Transformers | Embeddings | Semantic Search | Vector Search | Information Retrieval | Ranking | Text Classification | NER | Clustering | Model Fine‑tuning | Recall@K | MRR | NDCG

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Thompsons Hr Consulting • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Principal Data Scientist
Principal Data Scientist

Fractal Analytics Ltd. • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Optum India • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Senior AI Engineer
Senior AI Engineer

Genzeon Technology Solutions • Maharashtra

On-site
INR 3,500,000 - 6,000,000
Data Engineer
Data Engineer

AiLogic Neural Network Pvt Ltd • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Data Scientist
Data Scientist

Enterprise Minds, Inc • India

On-site
INR 4,200,000 - 7,000,000
Data Scientist
Data Scientist

AiLogic Neural Network Pvt Ltd • Hyderabad

On-site
INR 800,000 - 1,200,000
Data Scientist
Data Scientist

Celebal Technologies • Jaipur

On-site
INR 1,500,000 - 2,800,000
Junior AI/ML Engineer
Junior AI/ML Engineer

Kardee • Chennai District

On-site
INR 900,000 - 1,300,000
AI/ML Developer
AI/ML Developer

HubBroker ApS • Ahmedabad District

On-site
INR 1,200,000 - 2,400,000