An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Spotify is seeking a backend data engineer to own the data catalog, lineage and evolving schemas at scale. You will ensure data quality and governance as new storage technologies and data assets emerge.
Working with OpenLineage and Apache Iceberg standards, you’ll define naming, storage and metadata conventions that other teams rely on when publishing, querying and consuming data. This role emphasizes correctness and reliability, supporting platform teams with consistent access, traceability,
Every dataset at Spotify, from the data behind creator royalties to the signals powering recommendations, is registered, published and traced through systems our team builds. We're the source of truth for what data exists at Spotify, where it lives, who owns it, and how it flows from one pipeline to the next.
These systems sit on the critical path of the platform. Data teams use them every day to publish and find data, and other platform teams build on them for access, retention, governance and incident response. That means our work is judged on correctness and reliability, because every system built on top of ours is only as good as the metadata we provide.
This is a backend role focused on data management, working on the problems at the core of any large data platform: running a data catalog that stays accurate at scale, capturing lineage across batch and streaming workloads, managing schemas as they evolve, and guaranteeing that consumers only read complete data. We build on open standards like OpenLineage and Apache Iceberg, and define the internal standards for naming, storage and metadata that the rest of Spotify builds against. As the platform takes on new storage technologies and new kinds of data assets, those standards and the systems behind them have to evolve with it, and you'll help decide how.