Founding Software Engineer - Data Products

LH2 AI Labs

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

LH2 AI Labs is building the post-training infrastructure for frontier AI models and seeks a hands-on founding data/software engineer to own the data platform end to end. You will ingest from enterprise sources, normalize to a common shape, resolve identities, and assemble structured, sellable data units for AI labs.

This role requires crafting core architectural patterns, shipping fast in customer cloud environments, and guiding the team on data-model design and reliable pipelines.

Qualifications

  • 6+ years of data/software engineering with production pipelines ingesting from multiple external APIs
  • Strong Python and SQL capabilities
  • ETL/ELT experience across SaaS platforms and APIs
  • OAuth, pagination, rate limits, incremental/delta sync, backfills and idempotency
  • Data models, normalization layers, and entity resolution across sources
  • Experience with Snowflake/BigQuery/Databricks/Redshift and orchestration tools
  • Cloud deployments, preferably in customer VPC/private-cloud environments

Responsibilities

  • Build and operate production connectors across SaaS/API sources (Slack, Jira, GitHub, etc.)
  • Implement reliable incremental sync, pagination, retries, and backfills
  • Design the medallion data model and transformation layers
  • Develop tiered entity resolution with audit trails and confidence signals
  • Assemble structured, citation-backed data units for products
  • Handle schema evolution, deduplication, and per-tenant differences via config
  • Create idempotent, replayable, observable pipelines with versioned outputs
  • Establish data-engineering standards for the team

Skills

Python
SQL
ETL/ELT
Data pipelines
Architectural decisions
Cloud data workloads

Tools

Snowflake
BigQuery
Databricks
Redshift
Airflow
Dagster
Prefect

Job description

Employment Type: Full-Time
Mandatory requirement: Graduate from a Tier 1 engineering institution such as IIT, BITS, NIT, IIIT, or equivalent.
About LH2 AI Labs:

Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.

We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.

For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on—there is limited new signal left there. That is where we come in.

Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.

About the role:

We're building the platform that turns raw enterprise data — Slack, Jira, GitHub, email, docs, tickets — into structured, machine-trainable products for AI labs. You'll own the data foundation end to end: how we ingest from many enterprise sources, normalize the mess into one common shape, resolve identities across systems, and assemble the structured "decision" units we sell.

This is a hands-on founding role on a small team. You'll make the core architectural calls, set the engineering patterns the rest of the team builds on, and get a real product shipping inside customer cloud environments fast.

Responsibilities:
  • Build and operate production connectors across multiple SaaS/API sources (starting Slack, Jira, GitHub; expanding to Google Workspace, Teams, Notion, Confluence, CRM)
  • Implement reliable incremental sync, pagination, retries, backfills, and checkpointing
  • Design the medallion data model (raw → normalized entities → sellable products) and the transformation layers between them
  • Build tiered entity resolution — authoritative and deterministic cross-system identity joins first, probabilistic matching later — with confidence and evidence retained
  • Assemble structured, citation-backed "task/decision" units from resolved data, seeded from concrete outcomes (merged PRs, closed tickets)
  • Handle schema evolution, deduplication, late/deleted data, and per-tenant workflow differences via config, not code forks
  • Build pipelines that are idempotent, replayable, observable, and reproducible (versioned output manifests)
  • Establish data-engineering standards and patterns for the team
  • Work closely with the Privacy/PII founding engineer so the pipeline hands off cleanly at the trust boundary
Must-have:
  • 6+ years of data/software engineering, including owning a production pipeline that ingested from multiple external APIs with incremental sync, retries, and schema drift (this is the non-negotiable – not "I've used a pipeline," but "I built and operated one")
  • Strong Python and SQL
  • Strong ETL/ELT experience across multiple SaaS platforms and APIs
  • Deep understanding of OAuth, pagination, rate limits, incremental/delta sync, backfills, idempotency, and failure recovery
  • Experience designing data models, normalization layers, and entity resolution across heterogeneous sources
  • Experience with a modern warehouse/lakehouse (Snowflake, BigQuery, Databricks, Redshift, or equivalent) and orchestration (Airflow, Dagster, Prefect, or equivalent)
  • Comfortable owning architectural decisions and setting patterns in a small founding team
  • Experience operating data workloads in cloud environments, ideally customer VPC / private-cloud deployments
Nice-to-have:
  • Connector platforms (Airbyte, Meltano, Fivetran) and building custom API connectors
  • Entity-resolution tooling (Splink, Zingg) and record-linkage fundamentals
  • Identity/directory/HR data sources as an ER seed
  • Data lineage/catalog tooling (OpenLineage or similar)
  • Handling large volumes of semi-structured/unstructured data
  • Familiarity with AI training-data workflows (SFT, RLHF, evals) and what labs actually buy
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Member of Technical Staff (RL Environments and Evaluations)
Founding Member of Technical Staff (RL Environments and Evaluations)

LH2 AI Labs • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Lead Engineer
Lead Engineer

Arcadia • Chennai District

Hybrid
INR 4,000,000 - 7,000,000
Competitive compensation
Hybrid model with remote-first policy
Flexible Leave Policy
+7
Data Engineer - Data Platform
Data Engineer - Data Platform

Firmable • Kolkata District

On-site
INR 900,000 - 1,500,000
Data Engineer - Data Platform
Data Engineer - Data Platform

Firmable • India

On-site
INR 1,200,000 - 1,800,000
Senior Data Engineer - Data Platform
Senior Data Engineer - Data Platform

Firmable • Kolkata District

On-site
INR 400,000 - 800,000
Full Stack Engineer (Python)
Full Stack Engineer (Python)

Crisil • Mumbai, Pune District, Gurugram District

On-site
INR 1,800,000 - 3,000,000
Founding Engineer (Fullstack)
Founding Engineer (Fullstack)

Wisemonk • Hyderabad

On-site
INR 9,059,000 - 14,236,000
Senior Data Engineer - Data Platform
Senior Data Engineer - Data Platform

Firmable • India

On-site
INR 4,000,000 - 7,000,000
Equity
Flexible work style
Career growth opportunity
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
Data & Integration Analyst
Data & Integration Analyst

SCALIS • Pune District

On-site
INR 1,500,000 - 2,300,000