Stand out for this role — generate a tailored resume and cover letter in about a minute.
AMBSIMS & Associates Corp seeks a backend data engineer to own production data pipelines that turn raw enterprise data into clean, de-identified datasets for AI labs. You’ll work in a small team with ownership and autonomy, building multi-stage pipelines in Python on AWS and ensuring correctness and reliability.
You’ll apply NLP/NER techniques for de-identification, develop internal dashboards to monitor pipeline health, and ship features quickly while maintaining high data quality.
Location: Dumbo, Brooklyn, NY (onsite, 5 days/week)
Type: Full-time
Compensation: $220K-$300K base salary
Relocation: Up to $10K
Visa: Open to visa transfers (including OPT and H-1B). Additional sponsorship may be considered for the right candidate.
Openings: 2
Client is the primary source of real, proprietary enterprise data for the world's leading frontier AI labs. They acquire enterprise data generated through collaboration, communication, and building. They transform it into de-identified datasets that remain useful, and license those datasets directly to frontier AI labs.
They've scaled from $0 to a multi-eight-figure run rate in a matter of months. They have about 14 people, with a lean engineering team of around five, backed by Floodgate, Afore Capital, Ludlow Ventures, and Hustle Fund. The client was built by the team behind Sunset, which scaled to an eight-figure run rate and helped hundreds of venture-backed startups wind down.
You’ll own the backend systems that turn raw enterprise data into clean, de-identified, high-value datasets. This is not a role about moving data between systems. The work is code applied to the data itself: multi-stage pipelines that process large volumes of messy, real-world data where correctness matters as much as throughput.
You’ll join a small team where each engineer has real ownership, the problems are ambiguous, and your work directly affects revenue.
Required
Nice to have