Applied ML Engineer

Zywa

New York (NY)

On-site

USD 140,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Founding-level equity

Job summary

Cozmo is building the AI operating system for property claims in New York. The role focuses on turning raw claim data into scalable models that predict line-item outcomes and guide automated decisioning.

You will design schemas, ingestion, and feature pipelines, then train and deploy predictive models to support live claims. You will work directly with CTOs and founders, shaping the data platform and expanding capabilities from field photos to transcripts, while maintaining a bias toward

Qualifications

  • Excellent data engineering with Python, SQL and robust pipelines that survive schema drift.
  • Ability to own infrastructure end-to-end and work with evolving data schemas.
  • Strong applied ML on tabular data and knowledge of gradient-boosted models.

Responsibilities

  • Build the claims data platform from scratch: ingestion, schema design, feature pipelines, storage and orchestration.
  • Ship the first estimation models to production and iterate against live claim outcomes.
  • Design labeling and evaluation strategy for negotiated labels, where the metric is critical.
  • Wire model outputs into the agents running live claims, coordinating with forward deployed teams.
  • Build tooling to let a small ML effort scale to large data assets (training, tracking, serving, monitoring).
  • Extend into industry-first layers: scope prediction from field data and transcripts, leakage detection and pricing signals.

Skills

Python
SQL
dbt
Spark
Data pipelines
ML on tabular data
Gradient boosting
Production ML

Tools

dbt
Spark
Airflow

Job description

About Cozmo

Cozmo is the AI operating system for property claims. When a pipe bursts in someone's home at 2am, our agents answer the call, capture the loss, enter the claim into Xactimate and Cotality, dispatch the right contractor against SLA, chase acceptances before breach and draft the carrier-ready estimate from field photos. We run both halves of a claims operation: everything the customer touches and everything that happens behind the desk.


Our customers are restoration franchisors, TPAs and adjusting firms whose boards have told them to become AI-native and who have no way to do it themselves. Our anchor is one of the largest restoration franchisors in the US.


The role

Every claim we process generates hundreds of structured data points: line items, quantities, unit prices, room geometry, equipment counts, what was submitted and what the carrier approved. Years of this history sit in reports and carrier portals that nobody has ever built models on. Your job is to turn that raw sprawl into a data asset and ship the first models on top of it.


The pipeline is the hard part and the moat. Reports are messy, schemas drift across franchise locations, carrier responses arrive in half a dozen formats and the labels are negotiated outcomes rather than ground truth. The engineer who wins this role treats that mess as the job: builds ingestion that survives it, designs the schema the whole company will stand on and then trains the models, starting with gradient-boosted tabular systems predicting line-item approval outcomes and going wherever the data leads. Scope prediction from photos and transcripts, leakage detection and pricing intelligence are all open and unbuilt.


You own the whole path from raw export to a prediction serving a live claim. No handoffs, no research team upstream, no data team downstream. You are both.


You report to the CTO and work directly with both founders.


What you'll do


  • Build the claims data platform from scratch: ingestion, schema design, feature pipelines, storage and orchestration that a growing team inherits

  • Ship the first estimation models to production and iterate against live claim outcomes

  • Design the labeling and evaluation strategy for data where the label is a negotiated number, because getting the metric wrong here is worse than getting the model wrong

  • Wire model outputs into the agents running live claims, working with our forward deployed team during real go-lives

  • Build the tooling that lets a one/two-person ML effort move like ten: training pipelines, experiment tracking, serving, monitoring

  • Extend into the industry-first layer as the foundation solidifies: scope prediction from field photos and call transcripts, supplement and leakage detection


What we look for


  • 2 to 6 years shipping ML systems to production, with the pipeline scars to prove it

  • Excellent data engineering: Python, SQL, dbt or Spark or the equivalent, pipelines that survive schema drift and malformed real-world exports, comfort owning infrastructure end to end

  • Strong applied ML on tabular and structured data. You know when gradient boosting beats a neural network and you reach for the boring model that wins

  • Rigor about labels and leakage. You have caught a model that looked great in offline eval and was learning the wrong thing, and you can tell us exactly how

  • Bias for production over polish: you would rather have a live model improving weekly than a perfect one in a notebook

  • Evidence we can inspect: systems you built, data platforms you owned, open source, a writeup of a hard pipeline problem you solved

  • Excited to get close to the domain: reading estimates, sitting with adjusters, learning why a water mitigation claim prices the way it does


Nice to have


  • Structured extraction from documents, photos or call transcripts

  • Insurance, pricing, risk or marketplace data

  • LLM engineering for extraction and agent tool use


How we work

We work almost six days a week, in person in New York, and go-live weeks take the seventh day too. We say this plainly because intensity is our structural advantage against incumbents with a hundred times our headcount, and because the engineers we want read this section and relax. This is the environment where nobody tells you to slow down.


**Compensation**

Competitive salary for your market plus founding-level equity, weighted toward equity because the value of this role compounds with the data asset you build.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer
Forward Deployed Engineer

Zywa • New York (NY)

On-site
USD 120,000 - 190,000
Founding-level equity
Relocation to NYC
Forward Deployed Engineer
Forward Deployed Engineer

Cozmo AI • New York (NY)

On-site
USD 90,000 - 150,000
Agent Product Manager
Agent Product Manager

Zywa • New York (NY)

On-site
USD 120,000 - 180,000
Founding equity
Senior ML Ops Engineer
Senior ML Ops Engineer

United States Digital Space LLC • New York (NY)

On-site
USD 180,000 - 240,000
Equity
Fully paid health coverage
Dental and vision
+7
AI Engineer, Decision Intelligence
AI Engineer, Decision Intelligence

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Solutions Engineer
Solutions Engineer

Chariot Claims • Northern (KY)

Hybrid
USD 110,000 - 150,000
Domain Expert, Insurance
Domain Expert, Insurance

Sycamore • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Founding Staff AI Engineer
Founding Staff AI Engineer

Qumis • Chicago (IL)

Hybrid
USD 120,000 - 150,000
Medical, dental, and vision insurance
Unlimited PTO
Hybrid work environment
Applied ML Engineer — Claims Data Platform
Applied ML Engineer — Claims Data Platform

Zywa • New York (NY)

On-site
USD 140,000 - 230,000
Founding-level equity
Founding AI Engineer
Founding AI Engineer

Worky • San Francisco (CA)

On-site
USD 225,000 - 255,000