Senior Data Engineer II, Product Engineering

Polygon.io, Inc

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

401(k) plan
Unlimited time off
Medical plans

Job summary

Massive is seeking a data-focused engineer to own the middle of our data platform, transforming raw market, reference, and alternative data into fast, well-shaped datasets. You will define the data model, build the transformation layer in Python/SQL, and shape the API surface used by customers.

The role blends data engineering and analytics engineering, handling ingestion, cleaning, modeling, and production deployment with a hands-on, multi-hat team environment.

Qualifications

  • Strong Python and SQL skills with advanced analytics experience.
  • Experience modeling analytical data and choosing normalization vs denormalization.
  • Track record of accelerating table and query performance.
  • Experience with modern analytical engines or lakehouses.
  • Familiarity with columnar formats and object storage.

Responsibilities

  • Model the data, design entities, and document the schema rationale.
  • Build and own the transformation layer in Python and SQL for cleaning, normalization, validation, and reconciliation.
  • Ingest from external APIs, handle rate limits, pagination, and backfills with schema drift detection.
  • Preserve the source data while layering cleaned, curated data on top.
  • Make data systems fast via partitioning, sorting, and appropriate indexing.
  • Collaborate with Product to design API surfaces for datasets and queries.
  • Partner with Data Science/AI Engineering on datasets and features used by models.
  • Set data quality metrics for ingest validation, drift detection, freshness, and completeness.

Skills

Python
SQL
Data modeling
Performance optimization
Analytical data

Tools

DataFusion
DuckDB
Databricks
Snowflake
ClickHouse
Trino
BigQuery

Job description

Massive is building the financial market data platform developers actually want to build on. This role owns the middle of it: taking raw market, reference, and alternative data and turning it into modeled, well-shaped, fast datasets while also designing the API surface customers query them through.

Massive is building the financial market data platform developers actually want to build on. This role owns the middle of it: taking raw market, reference, and alternative data and turning it into modeled, well-shaped, fast datasets while also designing the API surface customers query them through.

This is a deliberately full-stack data role, spanning what many companies split into "data engineer" and "analytics engineer." You'll do some ingestion - mostly consuming third-party APIs, which means rate limits, pagination, incremental loads, backfills, and schema drift - but the center of gravity is everything after the data lands: cleaning it, modeling it, and deciding how it should be physically laid out so that queries stay fast as volume grows.

Most of this work is zero-to-one. You'll be handed a dataset and a direction rather than a spec, and you'll be expected to decompose it, make the calls, and drive it to something in production that other people depend on. The team is small enough that you'll wear several hats in a given month - modeling on Monday, chasing a performance regression on Wednesday, debating about an API contract on Friday.

We're looking for someone genuinely interested in the data itself: what the fields mean, how entities relate, where the edge cases live, and how all of that maps to the question a customer is actually trying to answer.

Responsibilities
  • Model the data. Design the entities, relationships, and grains that turn raw feeds into datasets people can reason about - and document why the model is shaped the way it is.
  • Build and own the transformation layer in Python and SQL: cleaning, normalization, validation, and reconciliation across overlapping sources.
  • Ingest from external APIs and feeds, handling rate limiting, pagination, incremental and backfill loads, schema-drift detection, and replay.
  • Preserve the source. Treat raw provider data as immutable and complete, and build cleaned and curated layers on top of it
  • Make it fast. Choose partitioning schemes, sort orders, file sizes, clustering, and indexes. Profile with EXPLAIN / query plans and engine metrics, and fix the slow path instead of adding hardware.
  • Partner with Product to design the API surface customers use to consume these datasets - resource and query design, filtering and pagination semantics, contracts, and versioning.
  • Partner with Data Science and AI Engineering, who are among your primary internal customers, on the datasets and features their work depends on.
  • Set the quality bar: validation at ingest, anomaly and drift detection in production, freshness and completeness expectations, and clear documentation of assumptions and known limitations.
Skills & Qualifications
  • Strong Python and SQL. Both, daily, and your SQL goes well past joins and aggregates - window functions, CTEs, and an instinct for what the planner is going to do with it.
  • Real experience modeling analytical data. You can walk us through a schema you designed, the tradeoffs you made, and what you'd change today. You know when to normalize and when to denormalize, and you can explain why.
  • A track record of making tables and queries faster - partitioning, file layout and sort order, indexing or clustering, and reading query plans to find the actual bottleneck rather than guessing.
  • Hands-on experience with at least one modern analytical engine or lakehouse: DataFusion, DuckDB, Databricks, Snowflake, ClickHouse, Trino, BigQuery, or similar.
  • Depth, not just exposure, in columnar formats and object storage. Parquet, Iceberg or Delta, S3-compatible storage - and you can explain why a columnar format wins for a given access pattern, not just that it does.
  • You drive. You take an ambiguous problem, break it down, pick a path, and get to something working - then say clearly what you'd fix next. You know when to ship a workaround versus fix the root cause.
  • You explain your thinking well. A lot of this job is making a modeling or storage decision legible to a researcher, an AI engineer, or the next person to touch the pipeline. If you can make a hard concept feel simple, that counts here.
  • Fluency with AI coding tools as a daily companion. We expect you to use Claude Code, Cursor, or equivalents to move faster on parsing, scaffolding, refactors, and boilerplate - and to steer them deliberately: good context, verification against the actual data, and rejecting output that's confidently wrong.
  • Curiosity about the domain. Prior financial data experience is a real advantage, but we would rather hire someone who asks sharp questions about a dataset than someone who has already seen this exact one.
  • Rust is preferred. We reach for it where performance matters, including work in and around DataFusion.
  • Arrow and Arrow Flight SQL for moving and serving result sets efficiently.
  • Experience working alongside data science or ML teams - feature pipelines, dataset versioning, evaluation sets.
  • Financial or market data is preferred: equities, options, corporate actions, reference data, fundamentals, or alternative data.
  • dbt, SQLMesh, or similar transformation and semantic-layer tooling.
  • Extracting structured data from irregular sources (XBRL, HTML, PDF).
About Massive

At Massive, we’re on a mission to modernize Wall Street by empowering developers with the tools to shape the future of finance. We’re reimagining financial market data for the 21st century, removing barriers, simplifying access, and creating frictionless, forward-thinking technologies.

Join us and become part of a passionate team that consistently sets new industry standards, creating a profound impact on the world of finance and technology, and leveling the playing field by providing fair access for all.

Massive is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. We strictly prohibit and do not tolerate harassment or discrimination against employees, applicants, or any other covered persons because of race, color, religion, creed, national origin or ancestry, ethnicity, sex, gender, gender identity, age, physical or mental disability, citizenship, sexual orientation, past, current or prospective service in the uniformed services. To request a reasonable accommodation, please email careers@massive.com .

Benefits for full time offers from Massive include, but are not limited to, comprehensive medical plans, 401(k), and unlimited time off. When determining a candidate’s compensation, we consider a number of factors including skillset, experience, job scope, and current market data.

Before submitting your application, please take a minute to review our Applicant Privacy Notice , which describes the data we collect, why we collect it, and how we use it. By submitting your application, you consent to our Applicant Privacy Notice .

Modernizing Wall St.

Reimagining financial market data for the 21st century.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer II, Product Engineering
Senior Data Engineer II, Product Engineering

Polygon.io • United States

On-site
USD 150,000 - 230,000
Senior Data Engineer II, Product Engineering
Senior Data Engineer II, Product Engineering

Massive • Atlanta (GA)

On-site
USD 120,000 - 180,000
Data Acquisition Engineer
Data Acquisition Engineer

Polygon.io, Inc • Northern (KY)

Hybrid
USD 110,000 - 170,000
Medical plans
401(k)
Unlimited time off
Technical Account Manager
Technical Account Manager

Polygon.io, Inc • Northern (KY)

Hybrid
USD 90,000 - 140,000
Quant Research Engineer, Derived Data Products
Quant Research Engineer, Derived Data Products

Polygon.io, Inc • Northern (KY)

Hybrid
USD 130,000 - 190,000
Medical plans
401(k)
Unlimited time off
Growth Specialist
Growth Specialist

Polygon.io, Inc • Northern (KY)

Hybrid
USD 90,000 - 120,000
Comprehensive medical plans
401(k)
Unlimited time off
Developer Support Engineer
Developer Support Engineer

Polygon.io, Inc • Northern (KY)

Hybrid
USD 120,000 - 170,000
Medical plans
401(k)
Unlimited time off
Principal Financial Data Ontologist & Taxonomy Architect
Principal Financial Data Ontologist & Taxonomy Architect

Polygon.io, Inc • Northern (KY)

Hybrid
USD 140,000 - 190,000
Comprehensive medical plans
401(k)
Unlimited time off
Technical Account Manager
Technical Account Manager

Polygon.io • United States

Hybrid
USD 80,000 - 110,000
Rust Engineer - Foundational Platform
Rust Engineer - Foundational Platform

Polygon.io, Inc • Northern (KY)

Hybrid
USD 120,000 - 160,000
Comprehensive medical plans
401(k)
Unlimited time off