Founding ML Engineer (India)

Crustdata (YC F24)

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Crustdata (YC F24) is seeking a Founding ML Engineer in Bengaluru, India. The role involves owning the research and engineering of core ML systems that process multilingual, web-scale data. Candidates should have over 3 years of experience in ML model development, particularly in NLP and entity resolution, with strong skills in Python and PyTorch. You'll be responsible for building and shipping models from research to production, making data searchable and enriched for various applications.

Qualifications

  • 3+ years building and shipping ML models in production.
  • Experience with transformer architectures and model fine-tuning.
  • Comfortable with contrastive, metric, and representation learning.
  • Expertise in using LLMs for structured extraction and data generation.

Responsibilities

  • Own ML systems turning multilingual data into structured intelligence.
  • Resolve data from multiple sources automatically.
  • Infer organizational structures from raw data.
  • Classify job changes and map raw titles.

Skills

Python
PyTorch
NLP
Information Retrieval
Entity Resolution
Text Classification

Job description

About The Role

Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification.

We're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers — and we went from 0 to $7M ARR in our first 12 months. Now we need someone who can push the boundaries of what our ML systems can do.

We're hiring a Founding ML Engineer to own the research and engineering behind our core intelligence layer. Our platform indexes hundreds of millions of professional profiles and company records from across the web. Making that data searchable, matchable, and enriched is an ML problem at its core.

This is not an MLOps role. You will be researching, training, and shipping models - from paper to prototype to production.

Who you are
  • 3+ years building and shipping ML models in production — NLP, information retrieval, or entity resolution
  • Strong with transformer architectures — you've trained and fine-tuned encoder models, not just called APIs
  • You know how to build and evaluate retrieval systems, classifiers, and embedding models
  • Comfortable with contrastive learning, metric learning, and representation learning
  • Experience using LLMs for structured extraction, classification, or data generation at scale
  • Strong Python and PyTorch
  • A true grinder — we work very hard
  • Founder mentality — someone who wants to be a founder in the future OR was a founder earlier
What you'll be doing

You’ll own the ML systems that turn messy, multilingual, web-scale data into structured intelligence. Some example problems:

  • A customer searches for 'RevOps professionals' — you need to return people titled 'Head of Revenue Department', 'Revenue Operations Manager', and 'VP Sales Operations' across English, French, and German
  • Three different data sources list what looks like three different companies — but it's actually one. You figure out how to resolve that automatically across millions of records
  • Given raw people data, infer the org chart — who reports to who, what the team structure looks like, how the engineering org differs from sales
  • Detect what technologies a company uses from unstructured signals scattered across the web
  • Classify whether a job change was a promotion, lateral move, demotion, or just a title edit — and do it for millions of transitions
  • Map raw job titles to canonical titles, seniority levels, and job functions — across dozens of languages and naming conventions
Nice to haves
  • Experience with entity resolution or record linkage at scale
  • Built taxonomy or ontology systems over messy real-world data
  • Background in multilingual NLP or cross-lingual transfer
  • Scaled LLM inference pipelines in production
  • Published research or open-source contributions in NLP/IR
  • Experience with distributed training on GPU clusters
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Engineer
AI/ML Engineer

Creuto Cloud Private Limited • India

On-site
INR 800,000 - 1,600,000
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)
AI Engineer (Python, GenAI/LLMs + ML Fundamentals)

Solutions By Text • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Optum India • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Senior Gen AI Engineer
Senior Gen AI Engineer

SciSpace • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Technical Lead – ML Engineer
Technical Lead – ML Engineer

Azilen Technologies • Ahmedabad District

On-site
INR 4,500,000 - 8,500,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

Dailoqa • Sector 10

On-site
INR 1,500,000 - 2,500,000
Founding Engineer - Data Products
Founding Engineer - Data Products

LH2 AI Labs • Bengaluru

On-site
INR 2,400,000 - 4,000,000
Lead Engineer - AI/ML
Lead Engineer - AI/ML

Mindfire Solutions • India

On-site
INR 2,000,000 - 3,000,000
AI/ML Engineer
AI/ML Engineer

Jash Data Sciences Pvt. Ltd. • Pune District

On-site
INR 800,000 - 1,200,000
Competitive salary
Learning opportunities
Exposure to latest AI technologies
AI Engineer
AI Engineer

Aziro • Chennai District, Bengaluru, Pune District

Hybrid
INR 1,200,000 - 1,800,000