Get more replies from employers
Send a job-specific resume in minutes.
Sunset is building systems to derive value from real enterprise data. We seek a hands-on Data Scientist focused on evaluation to define what high-quality, safe-to-deliver data means across de-identification, structure preservation, and semantic usefulness.
You will write Python and SQL, build evaluation corpora, and work with ML, Product Engineering, Data Engineering, Security, and Quality teams to drive measurable improvements.
About SunsetAt its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses.In 2025, we had a unique insight: the data every company generates each day through collaboration, communication, and building is some of the most valuable training data in the world. Public and synthetic data can only get frontier models so far, so the next generation of model progress depends on real, proprietary data grounded in how actual businesses operate. We are a primary source of it, partnering directly with the frontier AI labs building what comes next.
We have scaled from $0 to a multi-eight-figure run rate in a matter of monthsWe have raised from top-tier investors, including Floodgate, Afore, Ludlow, and Hustle FundWe are small enough that you will carry outsized responsibility and grow as quickly as the company doesYou will partner with and build for some of the fastest and most important companies in the worldYou will help build a massive, category-defining business from the ground floor
Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. That creates a difficult measurement problem. A system can improve aggregate F1 while missing a high-risk slice, remove more sensitive information while also destroying useful context, or pass one stage while defects escape somewhere else in the pipeline.As Sunset's first Data Scientist focused on evaluation, you will establish how we know whether that data is actually getting better. You will build the datasets, experiments, quality measures, and feedback loops that expose hidden failures, accelerate model and pipeline improvement, and give the team confidence in what it delivers.This is a hands-on, zero-to-one role at the intersection of data science, AI, and a real production system. You will write Python and SQL, construct evaluation corpora, study failure patterns, design comparisons, calibrate human and model-based judgments, and turn the result into a clear decision. The questions are scientifically difficult, but the output must be practical enough to change what the team builds and ships.You will work closely with Machine Learning, Product Engineering, Data Engineering, Security, Quality, domain experts, and the team making delivery decisions. Machine Learning Engineers own changing model behavior. You own the credibility of the evidence used to decide whether a model, pipeline, or delivery change actually made the data safer or more useful.