Staff Data Engineer

Harbor Compliance

United States

On-site

USD 172,000 - 215,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
Flexible paid time off
Parental leave
Fertility and adoption assistance
401(k)
Educational reimbursement

Job summary

Harbor Compliance is seeking a Staff Data Engineer to help build its first end-to-end data platform. You will design and implement real-time data pipelines and AI-ready embeddings from scratch, collaborating with the Sr. Director of Data Platform & Analytics.

In this hands-on role, you will own the stack from CDC events to vector databases, create dbt transformations, and partner across Finance, Marketing, and Ops to deliver reliable, low-latency data products for executive reporting.

Qualifications

  • 7+ years of hands-on data engineering experience in streaming systems.
  • Proven experience designing near real-time pipelines from scratch (Kafka, Kinesis, Flink, Debezium/CDC).
  • Hands-on production experience with vector databases and embeddings (e.g., Pinecone, Weaviate, pgvector, Milvus).
  • Advanced proficiency in SQL and Python.
  • Working knowledge of cloud warehouse/l lakehouse platforms (Snowflake, BigQuery, Databricks) and dbt.
  • Proven experience building end-to-end production data environments in early-stage teams.
  • Familiarity with B2B SaaS and recurring revenue data models.
  • Working knowledge of BI tools (Looker, Tableau, Power BI).
  • Ability to work independently in a lean, fast-moving environment.

Responsibilities

  • Design, build, and own near real-time data pipelines as the backbone of the platform's data flow.
  • Evaluate, implement, and maintain vector database infrastructure and embedding pipelines for AI-augmented use cases.
  • Design and build ELT/ETL pipelines ingesting data from our platform, CRM (HubSpot) and HRIS.
  • Architect the underlying warehouse/lakehouse as a system of record for streaming and AI infrastructure.
  • Build lightweight transformation layers (dbt) to deliver business-ready datasets.
  • Own pipeline reliability and observability across streaming and batch pipelines.
  • Enable self-service and AI-powered reporting with stakeholders on recurring reports.
  • Implement data governance with proper documentation and access controls.
  • Collaborate with Finance, Marketing, Customer Success, and Operations to translate data needs into reliable data products.
  • Leverage AI-augmented development tools to accelerate engineering and documentation.

Skills

Streaming data pipelines
Vector databases
SQL
Python
Cloud data warehouse
dbt
End-to-end data platform
BI tooling (Looker/Tableau/Power BI)
B2B SaaS data models
Event streaming (Kafka/Kinesis)

Tools

Kafka
Kinesis
Flink
Debezium/CDC
Snowflake
BigQuery
Databricks

Job description

Great Place to Work® Certified | USA

Harbor Compliance is a leading technology platform for entity compliance, helping more than 80,000 businesses and nonprofits manage licensing, tax registration, and legal entity requirements nationwide. Founded in 2012 and recognized repeatedly by the Inc. 5000 and Deloitte Technology Fast 500, we've grown through five strategic acquisitions ‑ and are now backed by a 2026 majority growth investment from Bregal Sagemount to accelerate product, AI, and customer experience. We're a passionate team making compliance simpler and smarter for every organization we serve.

Harbor Compliance is building its first end-to-end data platform

As Staff Data Engineer, you'll be a foundational technical hire, working closely with the Sr. Director of Data Platform & Analytics to design and build the real‑time, AI‑ready data infrastructure that powers this platform. This is a hands‑on role for someone who wants to build from the ground up, not maintain what already exists.

Key Responsibilities
  • Design, build, and own near real‑time data pipelines (CDC, streaming ingestion, event‑driven architectures) as the backbone of the platform's data flow
  • Evaluate, implement, and maintain vector database infrastructure and embedding pipelines to support AI‑augmented use cases (semantic search, retrieval‑augmented generation, AI agents acting on company data).
  • Design and build ELT/ETL pipelines ingesting data from our platform, financial platforms, CRM (HubSpot) and HRIS, feeding both real‑time and batch use cases.
  • Partner with the Sr. Director to architect the underlying warehouse/lakehouse as a supporting system of record ‑ the storage layer beneath the streaming and AI infrastructure.
  • Build lightweight transformation layers (e.g., dbt) as needed to enable our Analytics Engineering team translate raw data into business‑ready datasets aligned to core metrics like ARR, CAC, and churn.
  • Own pipeline reliability and observability ‑ monitoring, automated failure alerting, and lineage tracking across both streaming and batch pipelines.
  • Build the technical foundation for self‑service and AI‑powered reporting, partnering with BI, Product & Engineering stakeholders on recurring executive and departmental reports.
  • Implement data governance practices, including documentation standards and access controls.
  • Partner cross‑functionally with Finance, Marketing, Customer Success, and Operations to translate data needs into reliable, low‑latency data products.
  • Leverage AI‑augmented development workflows (e.g., Claude Code) to accelerate pipeline development and documentation.
Requirements
  • 7+ years of hands‑on data engineering experience, with meaningful depth in streaming/event‑driven systems, not just batch pipelines.
  • Proven experience designing and building near real‑time pipelines from scratch (e.g., Kafka, Kinesis, Flink, Debezium/CDC) in a production environment.
  • Hands‑on production experience with vector databases and embeddings (e.g., Zilliz, Pinecone, Weaviate, pgvector, Milvus) ‑ ideally having built this infrastructure from the ground up rather than just consuming a managed AI feature.
  • Advanced proficiency in SQL and Python.
  • Working knowledge of a cloud warehouse/lakehouse platform (Snowflake, BigQuery, or Databricks) and dbt ‑ you'll use these, but they're the storage/transform layer supporting the streaming and AI work, not the main focus.
  • Proven experience building or materially contributing to an end‑to‑end production data environment, ideally as an early or founding data hire.
  • Familiarity with B2B SaaS and recurring revenue data models (customer lifecycle, pipeline/conversion data).
  • Working knowledge of BI/reporting tools (e.g., Looker, Tableau, Power BI) as a downstream consumer of your data models.
  • Ability to work independently and drive multi‑stakeholder projects in a lean, scrappy, fast‑moving environment ‑ comfortable with ambiguity and building without a lot of existing infrastructure or process.
Skills And Knowledge
  • Strong command of event streaming and CDC tooling.
  • Hands‑on experience with vector databases (Pinecone, Weaviate, pgvector, Milvus, or similar), including embedding strategies and chunking approaches for retrieval use cases.
  • Working knowledge of ELT/ETL tools (e.g., Fivetran, Airbyte) for batch use cases.
  • Working knowledge of data observability/reliability practices ‑ automated alerting, lineage tracking.
  • Fluency with AI‑augmented development tools (e.g., Claude, Copilot) to accelerate engineering and documentation.
Accommodations

Harbor Compliance is committed to providing any reasonable accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter.

Compensation

Harbor Compliance’s base salary range for this role is listed below. Compensation at the time of offer is based on factors such as skill set, experience, qualifications, and work location. Salary is one part of Harbor Compliance’s total compensation package. Note that the salary range and benefits apply only to U.S.-based candidates.

  • Health benefits
  • Flexible paid time off
  • Parental leave
  • Fertility and adoption assistance
  • 401(k)
  • Educational reimbursement
Pay Transparency Policy Statement

Harbor Compliance will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by Harbor Compliance, or (c) consistent with Harbor Compliance’s legal duty to furnish information.

Equal Opportunity Statement

Harbor Compliance is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

Notice Regarding the Use of Selection Technology

Harbor Compliance uses third‑party automated tools and skills assessments (such as Testlify and Rippling) to help evaluate job applications and streamline our hiring workflow. These tools assist our recruitment team in reviewing qualification trends and scheduling, but all final employment decisions are made solely by humans on our Talent Success team, and all tools are subject to human oversight. Depending on your location, local laws may grant you specific disclosure rights or the option to request an alternative evaluation process. If you require a reasonable accommodation or wish to opt out of automated assessment steps due to regional regulations, please notify your recruiter. Opting out will not negatively impact your application.

The Pay Range For This Role Is

172,000 - 215,000 USD per year(US National)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Engineer
Staff Data Engineer

Harbor Compliance • Northern (KY)

Hybrid
USD 172,000 - 215,000
Business Development Representative
Business Development Representative

Harbor Compliance • United States

On-site
USD 48,000 - 60,000
Health benefits
Flexible paid time off
Parental leave
+1
Senior Product Manager
Senior Product Manager

Harbor Compliance • United States

On-site
USD 144,000 - 180,000
Manager, Partner Success
Manager, Partner Success

Harbor Compliance • Northern (KY)

Hybrid
USD 100,000 - 126,000
Health benefits
Flexible paid time off
Parental leave
+3
Senior Voice of the Customer Analyst
Senior Voice of the Customer Analyst

Harbor Compliance LLC • United States

Hybrid
USD 75,000 - 100,000
Health benefits
Flexible paid time off
401(k)
+1
Staff Data Engineer - Build Real-Time AI Data Platform
Staff Data Engineer - Build Real-Time AI Data Platform

Harbor-Compliance • United States

Remote
USD 172,000 - 215,000
Health benefits
Flexible PTO
Parental leave
+3
Manager, Partner Success
Manager, Partner Success

Harbor Compliance • United States

On-site
USD 100,000 - 126,000
Account Executive
Account Executive

Harbor Compliance • Northern (KY)

Hybrid
USD 60,000 - 90,000
Health benefits
Flexible PTO
Parental leave
+3
Principal Engineer
Principal Engineer

Harbor Lab • Kentucky

Hybrid
USD 190,000 - 240,000
30 days of paid annual leave
Private health insurance (family)
Hybrid or Remote work
+4
Senior Manager FP&A
Senior Manager FP&A

Harbor Compliance • Northern (KY)

Hybrid
USD 146,000 - 183,000
Health benefits
Flexible PTO
Parental leave
+2