- We are looking for a Senior Data Engineer to join the Data Platform team at DataSnipper
- Every decision DataSnipper makes about its products - which features land, which customers are getting value, what we bill for, what we fix next, which AI capabilities add the most value - runs through the data platform
- You will own the systems that make that possible: how usage events get captured across a growing set of products, how they become trustworthy models in Snowflake, and how every other team such as Customer Success, Product, and GTM teams get to the answers without waiting on us
- This is a hands‑on, high‑ownership role in a small team. We are a handful of people serving the whole company, so your judgment about what not to build matters as much as what you ship
- You will set the technical direction for ingestion and modeling, and you will be the person other engineering teams come to when they need to instrument something new
- The Data Platform team works across three areas, and this role sits closest to the first two:
- Data Platform - reliable, scalable infrastructure that gets the right data to the right place
- Internal Analytics - a self‑service platform so every team can be data‑informed without a ticket
- Customer‑facing Analytics - the dashboards and exports customers use to see the value they get from DataSnipper
- Concretely, you’d be walking into: billions of usage events flowing from our Excel Add‑in, web apps, and product backends through Azure Event Hubs into Snowflake; a dbt estate built on medallion principles and managed in dbt Cloud; Terraform‑managed Snowflake and Azure infrastructure; and a set of product teams shipping AI agents faster than we can instrument them
- You will also find real, named open problems rather than a tidy platform - event capture mid‑consolidation, multiple methods of user attribution, and a data quality layer that is designed but not yet built. We would rather tell you that up front
- Own the event ingestion architecture end to end - Azure Event Hub, Snowpipe, Fivetran, and our shared Python/TypeScript event client libraries
- Build and operate dbt transformation pipelines that stay reliable as volume, source count, and model complexity grow
- Define and enforce event contracts and schemas so product teams can instrument new features without silent breakage downstream
- Build reverse ETL and activation paths that push modeled data back into the tools the business works in - HubSpot properties and rollups, MongoDB, Postgres, and GTM reporting
- Evolve the core data models (event, user, license, company) that everything else depends on
- Own Snowflake performance and cost, and keep the platform’s tech debt, dependency, and compliance obligations (audit logging, vulnerability remediation, Vanta evidence) from accumulating
- Integrate and model new data sources across the business - product backends, MongoDB, HubSpot, billing, and third‑party tools
- Build the guardrails and tooling that let product teams create events, models, and dashboards themselves
- Contribute to the semantic / context layer so metrics have one agreed definition across BI tools, customer‑facing dashboards, and LLM and agent consumers
- Support the customer‑facing analytics surfaces (in‑product dashboards, standard and advanced data exports) with the aggregation and modeling work behind them
- Improve documentation and definitions to the point where analysts, stakeholders, and AI agents can self‑serve with confidence
- Partner with Product, Engineering, CS, and GTM to turn vague data requests into scoped, well‑defined work - and to push back when a request shouldn’t become a pipeline
Experience on a cloud platform at the infrastructure level (we’re on Azure; AWS/GCP transfers fine)Experience with event‑driven / streaming ingestion and the failure modes that come with it (schema drift, duplication, late data, backfills)7+ years in data engineering or a closely related backend/platform role, with a track record of owning a data platform area end to endExperience in a startup or scale‑up, especially as an early member of a data teamSolid data modeling fundamentals and the ability to defend a modeling decision to both engineers and business stakeholdersDeep SQL and strong Python, including query optimization and performance tuning on a cloud warehouseExcellent communication in English and genuine comfort working directly with non‑technical stakeholdersBias to action, sense of ownership, and the judgment to prioritize independently when demand exceeds capacityProduction experience with a cloud data warehouse (we use Snowflake) and a modern transformation framework (we use dbt)Experience with product analytics tooling (Mixpanel, RudderStack) and warehouse‑native BI (Netspring/Optimizely Analytics, Omni, Embeddable, or similar)Experience building data products for AI or agent consumption - semantic layers, metrics layers, MCP servers, or governed self‑service accessExperience with Terraform, Docker, and governance at scaleReverse ETL experience and familiarity with CRM data models (HubSpot, Salesforce) or customer success platformsExposure to B2B SaaS usage‑based pricing and entitlement data, or to audit/fintech