Own how a customer's data gets in, gets shaped, and stays correct as it moves through enrichment and into a release.
Everything the product promises depends on the data being right before anything clever happens to it. That is this role: the import path, the type inference, the identity column and the deduplication, the batch enrichment runs, and the immutable release snapshot that a live agent reads from.
The work is equal parts pipeline and correctness. A CSV with a decimal comma, a spreadsheet with merged cells, a file that claims to be 2MB and inflates to 2GB: each of these has already caused a real bug here, and each was fixed by making the pipeline refuse rather than guess.
You would also own the embedding and indexing path: what text represents a record, when it needs recomputing, and how the cache stays honest when a contract changes underneath it.
What you'd do
- Own the import path: CSV, TSV, JSONL and spreadsheets, with limits enforced on what is actually read rather than what a header claims
- Infer and confirm column types, identity columns and duplicates, so numeric sorting and dedup are correct downstream
- Build and tune batch enrichment: resumable runs, idempotent writes, and caches that never serve a stale answer
- Own the embedding pipeline: what gets embedded, when it is recomputed, and how it is stored
- Write the data tests: not just that the pipeline ran, but that what came out is what should have
What we need
- Three or more years building data pipelines that ran on a schedule and had someone depending on them
- Strong SQL and a working understanding of how Postgres behaves under load
- Python or TypeScript, and the judgement to know when a transformation belongs in the database
- You have handled genuinely messy real-world files and can describe what you did about encoding, nulls and duplicates
- You design for reruns: idempotency and resumability are instincts, not afterthoughts
Nice to have, not required
- Experience with vector stores, embeddings or search indexing
- Worked with a queue, a scheduler or a workflow engine in production
- Data quality or observability tooling you built rather than bought
How we work
- Small team, thin slices, shipped weekly. Nothing sits on a branch for a month.
- Reviews are real. Every feature ships with its tests, and a bug found in review is cheaper than one found by a customer.
- Safety-critical means the boring answer usually wins: fail closed, keep the receipt, don't guess.
- We only claim what ships. That applies to the product, the roadmap and the offer letter.
How hiring works
- 1 A 30-minute call about something you've built and what was hard about it.
- 2 A working session on a real problem from this codebase. You can drive, or we can pair.
- 3 A conversation about how you make decisions when the evidence is thin.
- 4 References, then an offer. We aim to answer within a week at every stage.