Founding Engineer - Data Products

LH2 AI Labs

Bengaluru

On-site

INR 2,400,000 - 4,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LH2 AI Labs is building the platform that turns raw enterprise data from Slack, Jira, GitHub, email, docs, and tickets into structured, machine-trainable products for AI labs. This hands-on founding role lets you own the data foundation end to end, from ingestion to normalized entities and sellable insights.

You will set engineering patterns, guide the architecture, and ship a real product inside customer cloud environments quickly.

Qualifications

  • 6+ years of data/software engineering with production pipelines
  • Strong Python and SQL
  • ETL/ELT across multiple SaaS platforms and APIs
  • Designing data models and entity resolution across sources
  • Experience with cloud environments and data warehouses
  • Comfortable making architectural decisions in a small founding team
  • Familiarity with OAuth, pagination, and data synchronization patterns

Responsibilities

  • Build and operate production connectors across SaaS/API sources
  • Implement reliable incremental sync, pagination, retries, and backfills
  • Design medallion data model (raw → normalized → sellable)
  • Build entity resolution with deterministic joins then probabilistic matching
  • Assemble structured, citation-backed units from resolved data
  • Handle schema evolution, deduplication, and per-tenant differences
  • Create idempotent, replayable, observable data pipelines
  • Establish data-engineering standards and patterns
  • Collaborate with Privacy/PII engineer for clean handoffs

Skills

Python
SQL
ETL/ELT
Data pipelines
OAuth
API integration
Data modeling
Cloud workloads
Snowflake/BigQuery

Job description

Employment Type: Full-Time

About LH2 AI Labs

Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.

We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.

For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on—there is limited new signal left there. That is where we come in.

Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.

About the role

We're building the platform that turns raw enterprise data — Slack, Jira, GitHub, email, docs, tickets — into structured, machine-trainable products for AI labs. You'll own the data foundation end to end: how we ingest from many enterprise sources, normalize the mess into one common shape, resolve identities across systems, and assemble the structured "decision" units we sell.

This is a hands-on founding role on a small team. You'll make the core architectural calls, set the engineering patterns the rest of the team builds on, and get a real product shipping inside customer cloud environments fast.

Responsibilities

  • Build and operate production connectors across multiple SaaS/API sources (starting Slack, Jira, GitHub; expanding to Google Workspace, Teams, Notion, Confluence, CRM)
  • Implement reliable incremental sync, pagination, retries, backfills, and checkpointing
  • Design the medallion data model (raw → normalized entities → sellable products) and the transformation layers between them
  • Build tiered entity resolution — authoritative and deterministic cross-system identity joins first, probabilistic matching later — with confidence and evidence retained
  • Assemble structured, citation-backed "task/decision" units from resolved data, seeded from concrete outcomes (merged PRs, closed tickets)
  • Handle schema evolution, deduplication, late/deleted data, and per-tenant workflow differences via config, not code forks
  • Build pipelines that are idempotent, replayable, observable, and reproducible (versioned output manifests)
  • Establish data-engineering standards and patterns for the team
  • Work closely with the Privacy/PII founding engineer so the pipeline hands off cleanly at the trust boundary

Must-have skills

  • 6+ years of data/software engineering, including owning a production pipeline that ingested from multiple external APIs with incremental sync, retries, and schema drift (this is the non-negotiable – not "I've used a pipeline," but "I built and operated one")
  • Strong Python and SQL
  • Strong ETL/ELT experience across multiple SaaS platforms and APIs
  • Deep understanding of OAuth, pagination, rate limits, incremental/delta sync, backfills, idempotency, and failure recovery
  • Experience designing data models, normalization layers, and entity resolution across heterogeneous sources
  • Experience with a modern warehouse/lakehouse (Snowflake, BigQuery, Databricks, Redshift, or equivalent) and orchestration (Airflow, Dagster, Prefect, or equivalent)
  • Comfortable owning architectural decisions and setting patterns in a small founding team
  • Experience operating data workloads in cloud environments, ideally customer VPC / private-cloud deployments

Nice-to-have skills

  • Connector platforms (Airbyte, Meltano, Fivetran) and building custom API connectors
  • Entity-resolution tooling (Splink, Zingg) and record-linkage fundamentals
  • Identity/directory/HR data sources as an ER seed
  • Data lineage/catalog tooling (OpenLineage or similar)
  • Handling large volumes of semi-structured/unstructured data
  • Familiarity with AI training-data workflows (SFT, RLHF, evals) and what labs actually buy
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer - Privacy & PII [ Data Products ]
Founding Engineer - Privacy & PII [ Data Products ]

LH2 AI Labs • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Founding AI Researcher
Founding AI Researcher

LH2 AI Labs • Bengaluru

On-site
INR 400,000 - 700,000
Delivery Lead-Data Engineer
Delivery Lead-Data Engineer

Acuity Analytics • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Data Science Engineer
Data Science Engineer

J&F • Chennai District

On-site
INR 1,200,000 - 1,800,000
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
Engineering Team Lead
Engineering Team Lead

Recro • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Founding ML Engineer (India)
Founding ML Engineer (India)

Crustdata (YC F24) • Bengaluru

On-site
INR 4,000,000 - 6,000,000
AI Engineer
AI Engineer

DocuSign, Inc. • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Software Engineer - Applied AI
Software Engineer - Applied AI

Hedgineer • Bengaluru

On-site
INR 1,800,000 - 4,000,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000