AI Data Engineer, Data Platform

Collective

San Francisco (CA)

Hybrid

USD 160,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid Work Model
Fresh Lunch
Transit reimbursement
Wellness reimbursement
Flexible PTO
Medical/dental/vision coverage
Parental leave
401k + equity
Annual in-person summit

Job summary

Collective is seeking a Data Engineer to own and scale the data platform powering analytics, reporting, and AI. You will design and maintain pipelines moving data from product, financial, and third-party systems into a BigQuery warehouse, and model it into clean, well-documented tables.

You will join the Data Engineering team and work with product engineers, analysts, and business stakeholders. This is a hands-on role focused on data quality, production ownership, and scalable foundations for

Qualifications

  • 5+ years in data/analytics engineering or similar, preferably in B2B SaaS or fintech.
  • Expert SQL and strong Python for building production pipelines.
  • Hands-on with BigQuery, dbt, Fivetran, and an orchestrator (Airflow/Dagster/Cloud Composer).
  • Deep data modeling in layered warehouse architecture with clear grain and naming.
  • Data quality, monitoring, alerting, and on-call in production.
  • Git-based workflows, CI/CD, and infrastructure-as-code; data infra as software.
  • Able to communicate trade-offs to non-technical stakeholders and drive alignment.

Responsibilities

  • Design, build, and maintain scalable data pipelines from apps, SaaS tools, and APIs into BigQuery.
  • Model data with dimensional design and layered architecture, ensuring clear documentation.
  • Own data quality, establish SLAs, and triage data incidents to root cause.
  • Tune performance and cost, including partitioning and clustering of warehouse data.
  • Set engineering standards for version control, code review, CI/CD, and IaC.
  • Govern and secure data with access controls and PII handling.
  • Collaborate with product, analysts, and stakeholders to enable self-serve reporting in Metabase.
  • Support AI and analytics use cases with a reliable semantic layer and metric definitions.

Skills

SQL
Python
BigQuery
dbt
Airflow
Data modeling
Data quality
Communication
Ownership
Git / CI-CD

Tools

BigQuery
Fivetran
dbt
Airflow
Dagster
Cloud Composer
Terraform
Datadog
Metabase

Job description

About Collective

Collective is on a mission to redefine the way businesses-of-one work. Our technology and team of trusted advisors help members achieve financial independence by taking care of everything from business incorporation to accounting, bookkeeping, tax services, and access to a thriving community, all in one integrated platform. We believe in empowering self-employed people to enjoy the same tax savings that big companies get, so they can focus on their passion, not paperwork.

Featured in Forbes, Business Insider, Yahoo, Bloomberg, Financial Times, TechCrunch, and more. We are backed by General Catalyst, Sound Ventures, QED Investors, Google’s Gradient Ventures, Expa, and other investors who have financed iconic companies like YouTube, Substack, Twitch, Box, Hims, Instacart, and Lyft.

About The Role

We are looking for a Data Engineer to own and scale the data platform that powers analytics, reporting, and AI across Collective. You will design, build, and maintain the pipelines that move data from our product, financial, and third-party systems into our BigQuery warehouse; model that data into clean, well-documented, reliable tables; and set the engineering standards that keep the platform trustworthy as the company grows.

You will join the Data Engineering team within Engineering and work closely with product engineers, analysts, and business stakeholders across Operations, Finance, and Go-to-Market. This is a hands-on role for someone who cares about data quality, takes ownership of production systems end-to-end, and wants their work to be the foundation the rest of the company builds on.

What You'll Do
  • Design and build data pipelines. Develop, deploy, and maintain scalable batch and event-driven pipelines that ingest data from application databases, SaaS tools, and external APIs into BigQuery using managed connectors (Fivetran), custom Python loaders, and orchestration tooling.
  • Model the data. Design and implement dimensional and analytical data models in dbt, following a layered architecture (raw, staging, marts) with clear grain, naming conventions, and documentation that analysts and downstream tools can rely on.
  • Own data quality and reliability. Implement testing, monitoring, alerting, and data contracts across the pipeline; define and meet freshness and accuracy SLAs; triage and resolve pipeline failures and data incidents to root cause.
  • Optimize performance and cost. Tune warehouse queries, partitioning, and clustering; manage BigQuery spend; and keep pipelines efficient as data volume grows.
  • Establish engineering standards. Drive best practices for version control, code review, CI/CD, and infrastructure-as-code across the data stack; document systems and runbooks so the platform is maintainable by the team.
  • Govern and secure data. Implement access controls, PII handling, and data retention practices appropriate for a financial services company; partner with Security and Legal on compliance requirements.
  • Enable the business. Partner with product engineers on source schema design and change management, and with analysts and stakeholders to translate business questions into reliable datasets, metric definitions, and self-serve reporting in Metabase.
  • Support AI and analytics use cases. Maintain the semantic layer, metric definitions, and documentation that allow LLM-based tools and internal agents to query the warehouse accurately and consistently.
What You'll Bring
  • Experience: 5+ years of professional experience in data engineering, analytics engineering, or a closely related role, ideally at a B2B SaaS or fintech company.
  • SQL and Python: Expert-level SQL and strong Python skills for building pipelines, transformations, and tooling; comfortable writing tested, production-grade code.
  • Modern data stack: Hands-on production experience with a cloud data warehouse (BigQuery strongly preferred), dbt or equivalent transformation framework, managed ingestion tools (Fivetran or similar), and an orchestrator (Airflow, Dagster, Cloud Composer, or similar).
  • Data modeling: Deep understanding of dimensional modeling, layered warehouse architecture, and schema design, with strong opinions on grain, naming, and consistency.
  • Data quality and observability: Experience implementing testing frameworks, lineage, monitoring, and alerting for data pipelines, and operating them in production including on-call.
  • Engineering fundamentals: Fluency with git-based workflows, code review, CI/CD, and infrastructure-as-code; you treat data infrastructure as software.
  • Ownership: A track record of taking ambiguous, high-impact problems and delivering reliable systems end-to-end, with a focus on outcomes rather than just implementation.
  • Communication: Ability to explain technical trade-offs to non-technical stakeholders and drive alignment on data definitions across teams.
Nice To Have
  • Experience with streaming or event data (Pub/Sub, Kafka, or similar) and product analytics tooling (Amplitude or similar).
  • Experience with Terraform and Google Cloud Platform infrastructure.
  • Experience with observability platforms such as Datadog.
  • Exposure to financial, accounting, tax, or payroll data and the correctness requirements that come with it.
  • Experience building semantic layers or metric stores consumed by LLM-based tools, or supporting LLM evaluation programs.
  • AI-assisted development experience (Claude Code or similar).
What We Offer
  • Hybrid Work Model: Based in San Francisco with a balance of in-office and remote flexibility.
  • Fresh Lunch: Provided on in-office days.
  • Commuter Support: $150 monthly reimbursement for transit expenses.
  • Health & Wellness: $200 quarterly reimbursement to support your well-being.
  • Time Off: Flexible PTO plus 14 company holidays.
  • Comprehensive Coverage: 100% medical, dental, and vision for employees; 75% coverage for dependents.
  • Parental Leave: 16 weeks fully paid.
  • Retirement & Ownership: 401k plan plus an equity package.
  • Team Connection: Quarterly virtual events and an annual in-person summit.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Data Engineer, Data Platform
AI Data Engineer, Data Platform

Expa • San Francisco (CA)

Hybrid
USD 140,000 - 190,000
Hybrid work model in San Francisco
Fresh lunch provided
Commuter and wellbeing support
+2
Fullstack Software Engineer
Fullstack Software Engineer

Collective Hub Inc. • San Francisco (CA)

On-site
USD 100,000 - 150,000
Hybrid Work Model
Fresh Lunch
Commuter Support
+6
Senior Software Engineer
Senior Software Engineer

Collective • San Francisco (CA)

On-site
USD 120,000 - 180,000
Hybrid Work Model
Fresh Lunch
Commuter Support
+6
Fullstack Software Engineer
Fullstack Software Engineer

Collective • San Francisco (CA)

On-site
USD 100,000 - 150,000
Hybrid Work Model
Fresh Lunch on in-office days
Commuter Support
+6
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Collective Hub, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 230,000
Hybrid work
Fresh lunch
Transit reimbursement
+2
Product Lead
Product Lead

Expa • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Hybrid work model
Fresh lunch
Commuter support
+6
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Collective • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Hybrid work model in SF
Fresh lunch
Commuter support
+4
Product Lead
Product Lead

Collective • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Fresh lunch provided
Transit reimbursement
+6
Product Lead
Product Lead

Collective Hub, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Hybrid work in SF
Fresh lunch provided
Commuter support: $150 monthly
+7
Staff Security Engineer
Staff Security Engineer

Collective Hub Inc. • San Francisco (CA)

On-site
USD 140,000 - 180,000
Hybrid Work Model
Fresh Lunch
Commuter Support
+6