Founding ML Engineer

Crustdata (YC F24)

San Francisco (CA)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI technology firm in San Francisco seeks a Founding ML Engineer to lead the research and engineering of their core AI systems. You will convert messy, multilingual web-scale data into structured intelligence, resolve data discrepancies, and train ML models from conception to production. The ideal candidate has 3+ years of experience in building production ML models in NLP or information retrieval, demonstrating a founding mentality and proficiency in Python and PyTorch.

Qualifications

  • 3+ years building and shipping ML models in production, specifically in NLP, information retrieval, or entity resolution.
  • Strong experience with transformer architectures, having trained and fine-tuned encoder models.
  • Familiarity with building and evaluating retrieval systems, classifiers, and embedding models.

Responsibilities

  • Owning ML systems to process multilingual web-scale data into structured intelligence.
  • Returning appropriate professional titles based on customer search queries.
  • Resolving disparate data sources indicating the same company across millions of records.

Skills

Python
PyTorch
NLP
Information Retrieval
Entity Resolution
Text Classification

Job description

About The Role

Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification.

We're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers — and we went from 0 to $7M ARR in our first 12 months. Now we need someone who can push the boundaries of what our ML systems can do.

We're hiring a Founding ML Engineer to own the research and engineering behind our core intelligence layer. Our platform indexes hundreds of millions of professional profiles and company records from across the web. Making that data searchable, matchable, and enriched is an ML problem at its core.

This is not an MLOps role. You will be researching, training, and shipping models - from paper to prototype to production.

Who you are
  • 3+ years building and shipping ML models in production — NLP, information retrieval, or entity resolution
  • Strong with transformer architectures — you've trained and fine-tuned encoder models, not just called APIs
  • You know how to build and evaluate retrieval systems, classifiers, and embedding models
  • Comfortable with contrastive learning, metric learning, and representation learning
  • Experience using LLMs for structured extraction, classification, or data generation at scale
  • Strong Python and PyTorch
  • A true grinder — we work very hard
  • Founder mentality — someone who wants to be a founder in the future OR was a founder earlier
What you'll be doing
  • A customer searches for "RevOps professionals" — you need to return people titled "Head of Revenue Department," "Revenue Operations Manager," and "VP Sales Operations," across English, French, and German
  • Three different data sources list what looks like three different companies — but it's actually one. You figure out how to resolve that automatically across millions of records
  • Given raw people data, infer the org chart — who reports to whom, what the team structure looks like, how the engineering org differs from sales
  • Detect what technologies a company uses from unstructured signals scattered across the web
  • Classify whether a job change was a promotion, lateral move, demotion, or just a title edit — and do it for millions of transitions
  • Map raw job titles to canonical titles, seniority levels, and job functions — across dozens of languages and naming conventions
Nice to haves
  • Experience with entity resolution or record linkage at scale
  • Built taxonomy or ontology systems over messy real-world data
  • Background in multilingual NLP or cross-lingual transfer
  • Scaled LLM inference pipelines in productionPublished research or open-source contributions in NLP/IR
  • Experience with distributed training on GPU clusters
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding ML Engineer
Founding ML Engineer

Open Select • San Francisco (CA)

On-site
USD 150,000 - 300,000
Salary range $150K to $300K
Equity at an early stage
Direct path toward founding opportunities
AI / ML Engineer
AI / ML Engineer

Neuron Factory • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Engineer Intern | Summer 2026
ML Engineer Intern | Summer 2026

Crustdata (YC F24) • San Francisco (CA)

On-site
Competitive stipend
Housing stipend for relocation
Direct mentorship from founders
+1
Founding Machine Learning Engineer [33189]
Founding Machine Learning Engineer [33189]

Stealth Startup • San Francisco (CA)

On-site
USD 140,000 - 240,000
Founding Machine Learning Engineer [33116]
Founding Machine Learning Engineer [33116]

Stealth Startup • San Francisco (CA)

On-site
USD 120,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

Aionia Group • Mountain View (CA)

Hybrid
USD 120,000 - 160,000
Senior ML Engineer
Senior ML Engineer

Next Ventures • New York (NY)

On-site
USD 130,000 - 160,000
ML Engineer – AI-Powered Automation & Workflow Intelligence
ML Engineer – AI-Powered Automation & Workflow Intelligence

Blue-Signal-Search • San Francisco (CA)

On-site
USD 130,000 - 160,000
Competitive compensation package
Significant equity upside
Collaborative in-person work environment
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Machine Learning Engineer
Machine Learning Engineer

zaimler • San Mateo (CA)

On-site
USD 140,000 - 190,000