Principal Data Engineer

Gather AI

United States

Remote

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Gather AI is hiring a Principal Data Engineer to architect a data platform from the ground up, separating analytics from live production, and building a semantic metrics layer. You’ll ensure traceability, define metrics once, and support multiple product lines including drone imagery integration.

You will collaborate across engineering, product, and leadership to scale data infrastructure, drive data quality, and enable self-service analytics for internal teams and stakeholders.

Qualifications

  • Over 10 years in data engineering or closely related field.
  • Proven experience building greenfield data warehouse and OLTP→OLAP transition.
  • Deep expertise in multi-layer transformation architectures and reusable frameworks.

Responsibilities

  • Architect a greenfield, multi-layer data warehouse (raw, refined, serving).
  • Deliver a governed, self-service data-access layer for internal consumers (Product, CSM, Deployment/Operations, Leadership).
  • Build a semantic/metrics layer with consistent definitions across dashboards and products.
  • Own data-ingestion correctness at the boundary with the integration/backend team.
  • Stand up a data catalog and lineage layer (Purview or DataHub) for data ownership and traceability.
  • Prove end-to-end foundation on drone product, then generalize for other products.
  • Act as connective tissue between product and ML, linking structured data to drone imagery with traceability.

Skills

Data engineering
Data platform architecture
SQL
dbt
Azure
Governance
Multi-tenant
Streaming data
ETL/ELT
Cloud architecture

Education

Bachelor's degree in CS/Engineering

Tools

Purview
DataHub
dbt
Airflow
Azure

Job description

About Us

Are you ready to build the future of the supply chain? At Gather AI, we’re not just creating software; we’re pioneering a new era of warehouse intelligence. We’ve developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining “on-time, in full” delivery.

If you’re looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We’re leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time.

About the Team

You’ll join the Data Platform team at its inception, helping establish the foundation from day one. Today, production and analytical workloads share a single database, and every product team defines its own metrics. This team exists to fix that: designing the warehouse, the transformation layers, and the semantic model that every product and dashboard will build on going forward, in close partnership with product, engineering, and security.

About the Role

Most data platform roles ask you to extend something someone else has already built. This one starts with a blank canvas.

As Principal Data Engineer , you’ll architect Gather AI’s data foundation from the ground up: separating analytical workloads from live production traffic, building a semantic layer so metrics are defined once and stay consistent everywhere, and linking structured records to real drone imagery and video with full traceability. You’ll prove the model end to end on Gather’s drone product, then generalize it so every new product extends the foundation instead of rebuilding it, all while working as a Principal-level individual contributor with real influence across engineering, product, and leadership.

What You’ll Do
  • Architect a greenfield, multi-layer data warehouse (raw, refined, serving) that separates analytical workloads from production OLTP traffic.
  • Deliver a governed, self-service data-access layer for internal consumers first (Product, CSM, Deployment/Operations, and Leadership) as Phase 1, ahead of customer-facing conversational analytics.
  • Build a semantic and metrics layer so every metric, such as “scan accuracy by site,” is defined once in code and stays identical across every dashboard and product, making self-service safe from metric drift.
  • Own the quality bar: 99%+ availability SLA with freshness guarantees, 100% traceability, zero cross-tenant leakage, 99.5%+ pipeline success, and no data loss.
  • Design tenant isolation, per-tenant cost attribution, and schema and row-level RBAC to scale toward hundreds of tenants (300+ target), not today’s fleet size.
  • Own data-ingestion correctness at the boundary with the integration/backend team, covering data contracts, schema validation, and pipeline quality, so WMS data lands in the right place, shape, and time across WMS versions.
  • Stand up a data catalog and lineage layer (Purview as the Azure-native fit, DataHub as the open-source alternative) so every consumer can find data, see ownership, and trace lineage when a metric looks wrong.
  • Prove the foundation end to end on Gather’s drone product, then generalize it so each new product extends the model instead of rebuilding it
  • Act as the connective tissue between product and ML (3DCC, damage detection). Link structured records to unstructured drone imagery and video with full traceability, and stand up the data-infra readiness for feature stores and annotation pipelines on one trusted foundation.
What You’ll Need
  • 10+ years in data engineering, with 3+ years architecting data platforms for data products, analytics, or AI-driven products.
  • Proven experience building a greenfield data warehouse and leading an OLTP to OLAP transition, not just maintaining an existing one.
  • Deep expertise designing multi-layer transformation architectures and reusable frameworks that scale across multiple product areas.
  • Expert SQL and dbt, hands-on ELT and orchestration, and large-scale or streaming data experience.
  • Production experience on a major cloud (Azure preferred, AWS or GCP acceptable), plus infrastructure as code and CI/CD.
  • Track record with data quality, security, governance, and multi-tenancy in production environments.
  • Data transformation and modeling that turns raw multi-source data into refined, serving-ready datasets (raw to refined to serving).
  • Pipeline orchestration and workflow automation for scheduling, dependency management, and reliable execution across data flows.
  • Large-scale and distributed processing of high-volume batch data.
  • Real-time and streaming ingestion that captures and processes event data as it arrives.
  • Semantic and metrics-layer design that defines business metrics once and serves them consistently to every consumer.
  • Serving-layer optimization for fast, low-latency consumption through wide and flattened tables and pre-computed metrics.
  • Cloud data engineering and infrastructure automation that provisions, deploys, and operates the platform reproducibly (cloud-native, infrastructure as code, CI/CD).
  • Data quality, observability, and lineage that ensure trust, freshness, and end-to-end traceability.
  • Security, governance, and multi-tenancy including tenant isolation, access control, and resiliency.
  • Multimodal data integration that links structured records to unstructured image and video (drone captures) with traceability.
Ways of Working
  • Treats data as a product for internal consumers, not just a pipeline feeding dashboards.
  • Comfortable making long-lead architecture calls (platform, isolation model) with incomplete consensus.
  • Strong cross-functional collaborator, works closely with integration/backend, ML, product, customer success teams and internal analytics consumers.
Nice to Have
  • Experience modeling structured data linked to unstructured or blob data such as images, video, or sensor files
  • Experience with feature stores, annotation pipelines, or ML data infrastructure supporting computer vision products.
  • IoT, edge, or device-telemetry background
  • BI or presentation-layer and dashboard design experience
  • Warehousing, logistics, or supply-chain domain knowledge
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Data Engineer
Principal Data Engineer

Caraa • Indianapolis (IN)

On-site
USD 150,000 - 190,000
Member of Data Staff (Analytics Engineer)
Member of Data Staff (Analytics Engineer)

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Principal Software Development Engineer
Principal Software Development Engineer

Auger Inc. • Bellevue (WA), Northern (KY)

On-site
USD 190,000 - 250,000
Data Platform Engineer
Data Platform Engineer

The Phoenix Group • Stamford (CT)

On-site
USD 130,000 - 210,000
Data Engineer
Data Engineer

The Phoenix Group • Scottsdale (AZ)

On-site
USD 120,000 - 180,000
Principal Data Engineer
Principal Data Engineer

Worth AI • Miami (FL)

On-site
USD 130,000 - 170,000
Health Care Plan
Retirement Plan
Unlimited Paid Time Off
+2
Senior Data Engineer
Senior Data Engineer

Madison-Davis, LLC • Chicago (IL)

On-site
USD 130,000 - 180,000
Principal Data Engineer — Architect Greenfield Platform
Principal Data Engineer — Architect Greenfield Platform

Gather AI • United States

On-site
USD 180,000 - 270,000
Data Engineer
Data Engineer

Deltek • United States

Remote
USD 140,000 - 230,000