Get more replies from employers
Send a job-specific resume in minutes.
Product.ai in Santa Monica, Los Angeles is hiring a Senior Data Engineer to own the commerce data supply chain end-to-end, from ingestion to canonical records and freshness.
You will design and operate the agent validation fleet, set data-quality SLAs, and build production pipelines with Python and SQL, using Airflow or Dagster to orchestrate large-scale data flows.
Own every code and merchant fact from source to shelf, across hundreds of thousands of stores: what enters, how it is normalized, and how fast it goes stale.
Product.ai is the verified truth layer for shopping. When a person or an AI agent needs to know what is actually true about a purchase, we answer with proof. SimplyCodes is the first proof at scale, the code verification service that shows shoppers codes that actually work. It earns about $22 million a year at roughly 60% margins. We are 100% founder-owned, profitable, and bootstrapped since 2009. No outside investors, no board. Fewer than twenty operators, outbuilding companies 10x our size.
Every claim we publish stands on a supply chain of facts. Codes and merchant facts arrive by the hundreds of thousands from networks, feeds, and merchant surfaces we do not control, each in its own shape and wrong in its own way. They resolve into one canonical shelf: this code, this merchant, these terms. And they start dying the moment they land. Codes expire and terms change without notice. Freshness is the fuel of the engine, because a stale shelf is a wall of dead codes, and dead codes are what we exist to kill.
That supply chain has never had a single owner. The pipelines run and the revenue flows, but no one seat has ever decided what enters, what record it becomes, and how fast each fact is rechecked or retired. This is a founding seat, the whole intake-to-shelf path in one pair of hands, working directly with the founder. One layer is mid-transition. Row-level validation, the industry's last manual stronghold, is becoming an agent fleet here. LLM pipelines read and check at machine scale; humans own the verdicts. You inherit an agent-fleet design problem, not a headcount.
One boundary, drawn on purpose. Downstream of you, a robot fleet proves codes at real checkouts, and a sibling seat owns that proof. You own everything upstream, including the clock that decides when a fact must be proved again. They prove what you shelve; you decide what is worth proving. Two seats, one loop.
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
You reason about data systems in invariants and lifecycles. Handed a shelf of facts you have never seen, you first ask where each row came from, what would make it wrong, and when it dies. You see the pattern behind the pile, like the five sources behind one duplicate or the decay class behind one stale code. You notice when your model is wrong and update fast, and you write clearly, because on a team this small the written spec is the meeting.
You move between architecture and shipped code without ceremony, and a resolution design in the morning can be processing real rows by night. Agents are your production workforce; you direct them and verify what comes back. You can do this job by hand and prove it, and that mastery is what lets you trust, or reject, what an agent hands you. The expensive thing here is a redo cycle, never the compute.
You have built and run production data pipelines at real scale: ingestion from third parties you did not control, entity resolution or dedupe where a bad merge cost something, data-quality guarantees someone else depended on. That record can come from catalog and listings platforms, price intelligence, ad or affiliate data, search indexing, or knowledge graphs. All the same discipline. What counts is that your guarantees held. Python and SQL are daily tools; Airflow, Dagster, or Temporal are familiar ground; your batch-versus-streaming opinions come with incident stories attached. If you have run LLM pipelines against golden sets, better still. If not, you will learn that here fast. We care about the artifact and the reasoning far more than where you did it, and there is no degree to check.
This seat is wrong if you guard one lane and call the rest someone else's department; the supply chain runs from raw source to public shelf, and you own all of it. It is wrong if you need a finished spec and a groomed queue before you can move, or a platform team underneath you to feel senior. It is wrong if you pick tools for the resume line rather than for what the pipeline needs tonight. And it is wrong if you would ship what an agent handed you without being able to say why it is right, or let the fleet grade its own homework. You will be happiest here if your idea of craft is a shelf that is never silently wrong, and a supply chain you can defend row by row.
We don't run traditional engineering interviews. We evaluate demonstrated performance on work-relevant tasks, in four steps.
We hire on the work and the reasoning, not the pedigree.
Total first-year comp: $400,000 – $500,000 — base, plus performance-based ownership and profit-share programs. Base: $280,000 – $330,000, top of market for senior data engineering.
Eligibility for the company's ownership and profit-share programs — grants are performance-based, with terms discussed at the offer stage; 100% family premium coverage; an AI tooling budget steered by return, never capped. The model is built to mint partners.
Based in Santa Monica, Los Angeles — in person, five days a week. The rooms are real rooms. Relocation support available for the right builder.