Data Engineer - Data Platform

Firmable

India

On-site

INR 1,200,000 - 1,800,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Firmable is seeking a Data Engineer — Data Platform to build pipelines turning billions of records into clean, modelled datasets for product, analytics, and AI systems.

You’ll own LLM-based extraction, enrichment, validation in the pipeline, plus reliability, cost, and data trust across Snowflake, dbt, and Airflow. Strong Python/SQL and end-to-end pipeline ownership are essential.

Qualifications

  • 3+ years in data engineering or data infrastructure with end-to-end pipelines.
  • Strong Python and SQL with large-scale performance.
  • Shipped LLMs inside data pipelines with structured outputs and a labelled eval set.
  • Harness engineering experience for eval/test harnesses.
  • Sharp judgement on rules vs. LLMs with regex or dbt tests.
  • Production dbt and Airflow with modular, tested models.
  • Cloud warehouse experience, Snowflake preferred.
  • AI coding tools used daily with real shipped work.
  • Ownership and systems thinking across dependencies.

Responsibilities

  • LLM-in-the-pipeline systems for extraction, enrichment, entity resolution and validation.
  • Eval harnesses with labelled sets and regression suites gating model changes.
  • Rules vs. LLMs: deterministic checks where possible and semantic checks where needed.
  • Cost and drift management with token budgets and model routing.
  • Observability: log LLM calls with prompts, model, cost, latency and decisions.
  • Core pipelines and warehouse: Airflow, dbt models, Snowflake performance and cost.
  • Matching and deduplication across 13 markets.

Skills

Python
SQL
Pipelines ownership
Eval harnesses
LLM integration
dbt
Airflow
Cost optimization
Data quality
Model integration

Tools

dbt
Airflow
Snowflake
AWS
OpenTelemetry
Claude Code
Cursor

Job description

Firmable is the market-leading B2B sales intelligence platform in Asia Pacific — and we're scaling that success globally at pace. Backed by leading investors and 2,000+ customers strong, we exist to give sales teams an unfair advantage: the deepest company and people data of any platform, enriched with real-time signals, served at the right moment by intelligent agents.

Our data is the product. Building it now means building with LLMs in the pipeline, and knowing exactly when to trust them.

The Role

As Data Engineer — Data Platform, you'll build the pipelines that turn billions of raw records into the clean, modelled datasets our product, analytics, and AI systems depend on. Some of that is classic data engineering: SQL, dbt, orchestration, warehouse performance. A growing share is not: LLM-based extraction, enrichment, and validation running inside the pipeline at scale.

The hard part isn't calling a model. It's making a probabilistic component behave like a reliable one: labelled eval sets, precision and recall you can measure, versioned prompts, cost ceilings, and the harnesses that catch regressions when a vendor silently changes the model under you.

This is a hands-on engineering role with real ownership. You'll own reliability, cost, and whether downstream consumers can trust the data, whether it came from a SQL join or a language model.

What You'll Own
  • LLM-in-the-pipeline systems — extraction, enrichment, entity resolution, and semantic validation steps that run over millions of records a day with structured outputs, retries, and human-review fallbacks

  • Eval harnesses — labelled sets, scorers, and regression suites that gate every prompt or model change; precision/recall tracked per check, not vibes

  • Rules vs. LLMs — deterministic checks (dbt tests, data contracts, SQL) wherever structure allows; LLMs where semantic judgement is needed; the discipline to know which is which

  • Cost and drift — token budgets per pipeline, model routing (cheap models for classification, stronger ones for hard cases), drift detection on vendor updates

  • Observability — every LLM call logged with prompt version, model, cost, latency, and decision, alongside standard pipeline alerting that surfaces problems before they cascade

  • Core pipelines and warehouse — Airflow orchestration, dbt models across staging to mart, Snowflake performance and cost over billions of rows, AWS infrastructure as code

  • Matching and deduplication — embeddings and retrieval patterns for company and people entity resolution across 13 markets

What We're Looking For

Must Haves

  • 3+ years in data engineering or data infrastructure, with production pipelines you've owned end to end and 2+ years experience with Snowflake

  • Strong Python and SQL — production-grade, performance-aware, comfortable at very large scale

  • Shipped LLMs inside data pipelines — extraction, enrichment, or validation in production, with structured outputs and a labelled eval set that tells you where the model gets it wrong

  • Harness engineering experience — you've built or maintained eval or test harnesses for LLM outputs, and you can talk through what they caught

  • Sharp judgement on rules vs. LLMs — you reach for a regex or a dbt test first and can defend the call either way

  • Production dbt and Airflow — modular, tested models; DAGs that recover gracefully

  • Cloud warehouse experience — Snowflake preferred; schema design, query optimisation, cost management

  • AI coding tools are how you work — Claude Code, Cursor, or equivalent, daily, with real shipped work to show for it

  • Ownership and systems thinking — you weigh upstream dependencies and downstream impact before changing anything

Highly Valued

  • Eval frameworks and LLM tracing (Logfire, OpenTelemetry)

  • Embeddings, vector search, or fuzzy matching for entity resolution at scale

  • Fine-tuning or distilling small models to replace expensive LLM calls

  • AWS at scale (S3, Lambda, Glue, ECS, RDS) and PostgreSQL

  • Spark or PySpark; streaming (Kafka, Kinesis, Snowpipe Streaming)

  • Data privacy and compliance (GDPR, SOC2, CCPA)

How We Build

Firmable is an AI-native organisation. AI coding tools, automated testing, and AI-assisted review are how we work by default. Every LLM check ships with a labelled eval set, measured precision/recall, and a prompt version you can roll back. Every LLM call in production is logged from day one; retrofitting later is not the plan.

We run lean and ship fast — small senior teams, no layers, minimal process, weekly releases moving toward daily. Teams own their stack end to end. There are no fixed hours and no handholding. If you're not already working this way, this role isn't right for you.

Why This Role
  • LLMs as production infrastructure — not a demo, not a notebook; models making millions of decisions a day on data customers pay for

  • Greenfield harnesses — eval coverage, drift detection, and cost controls for in-pipeline LLMs are largely unbuilt; you'll ship them

  • Scale that matters — billions of rows, 13 markets, and a dataset nobody else has

  • Small team, massive leverage — your pipelines reach every Firmable customer, every day

  • Competitive base + meaningful equity — a share in the upside we're building toward

Firmable is an equal opportunity employer. We believe diverse teams build better products.

Ready to build the AI-native data platform behind the world's smartest B2B sales intelligence platform?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - Data Platform
Data Engineer - Data Platform

Firmable • Kolkata District

On-site
INR 900,000 - 1,500,000
Senior Data Engineer - Data Platform
Senior Data Engineer - Data Platform

Firmable • Kolkata District

On-site
INR 400,000 - 800,000
Senior Data Engineer - Data Platform
Senior Data Engineer - Data Platform

Firmable • India

On-site
INR 4,000,000 - 7,000,000
Equity
Flexible work style
Career growth opportunity
Data Engineering Lead - Data Quality Systems
Data Engineering Lead - Data Quality Systems

Firmable • India

On-site
INR 3,500,000 - 6,000,000
Software Engineering Lead - Data Services
Software Engineering Lead - Data Services

Firmable • India

On-site
INR 3,000,000 - 6,000,000
Founding Software Engineer - Data Products
Founding Software Engineer - Data Products

LH2 AI Labs • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Engineer – Data
Senior Engineer – Data

Orbital • Hyderabad

On-site
INR 2,000,000 - 4,000,000
Data Scientist
Data Scientist

Emergent • Bengaluru

On-site
INR 2,600,000 - 4,200,000
Daily Meals: Lunch and Dinner provided
Family Insurance: 3 Lakhs coverage for
Unlimited Paid Time Off
+1
Data Scientist New Bengaluru
Data Scientist New Bengaluru

Emergentagent • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Daily Meals
Family Insurance
Unlimited Paid Time Off
+1
Senior AI Devops Engineer
Senior AI Devops Engineer

Demandbase • Hyderabad

On-site
INR 4,500,000 - 7,000,000
LLM gateway experience
LLM observability tooling
Eval frameworks
+1