Data Engineer

Drive Capital

Toronto

On-site

CAD 80,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Professional development support
Dynamic work environment

Job summary

Drive Capital is seeking a skilled Data Engineer in Toronto to build AI platform data pipelines using Databricks and Spark. The ideal candidate has over 4 years of data engineering experience and strong skills in Python, SQL, and event streaming systems. Responsibilities include designing ETL pipelines, managing data quality, and integrating third-party data sources. You’ll work in a dynamic startup environment with competitive salary and opportunities for professional growth. This is a hybrid role with 4 days per week in the office.

Qualifications

  • Strong proficiency in Python and SQL for data transformation.
  • Production experience with Spark (PySpark) and distributed data processing.
  • Solid understanding of data modeling patterns.

Responsibilities

  • Design, build, and maintain ETL pipelines using Databricks and Spark.
  • Build event-driven data flows using GCP Pub/Sub.
  • Prepare and manage datasets for LLM fine-tuning.

Skills

Python
SQL
Spark (PySpark)
Cloud data platforms
Event streaming systems

Education

4+ years data engineering experience

Tools

Databricks
BigQuery
GCP Pub/Sub
dbt

Job description

About Us

The rules of discovery have changed. The search bar is becoming a conversation, and brands need a new playbook to win. That's where Yolando comes in. We are the command center for the AI era, helping marketers move from simple visibility to true velocity. Backed by $12M in funding (including Drive Capital and MaRS Discovery District), our tight-knit team of 15 is building the engine that defines how brands get found, cited, and recommended by AI. We aren't just building a roadmap; we're building the standard for Generative Engine Optimization.

Role Overview

We are seeking a skilled Data Engineer to build the backbone of our AI platforms, Yolando and BirdseyePost. You will design and maintain sophisticated ETL pipelines using Databricks and Spark, ensuring the reliable flow of data that powers our insights and ML models. You will implement Bronze-Silver-Gold medallion architectures and build event-driven flows to process streaming data for real-time analytics. In this role, you will prepare datasets for LLM fine-tuning and drive the integration of third‑party sources to enable data‑driven decision‑making at scale.

Key Responsibilities
  • Build and Optimize Data Pipelines: Design, build, and maintain ETL pipelines using Databricks and Spark for processing customer data, campaign analytics, and AI model inputs. Implement Bronze-Silver-Gold medallion architectures for reliable data transformation.

  • Enable Real‑Time Data Processing: Build event‑driven data flows using GCP Pub/Sub and Protocol Buffers. Process streaming data for real‑time analytics, attribution tracking, and AI system inputs.

  • Power AI and ML Systems: Prepare and manage datasets for LLM fine‑tuning, embedding generation, and recommendation systems. Build pipelines that feed vector databases (pgvector) with processed embeddings for semantic search.

  • Integrate Third‑Party Data Sources: Build reliable ingestion pipelines for platforms like Klaviyo, Shopify, and marketing APIs. Handle incremental loads, schema evolution, and data quality validation.

  • Drive Analytics and Attribution: Implement attribution models, customer lifetime value (CLV) calculations, and campaign performance analytics. Build data models that power dashboards and enable data‑driven decision making.

  • Ensure Data Quality and Reliability: Implement data validation, monitoring, and alerting for pipeline health. Build idempotent, retry‑safe pipelines that handle failures gracefully.

What We're Looking For
  • 4+ years data engineering experience.

  • Strong proficiency in Python and SQL for data transformation.

  • Production experience with Spark (PySpark) and distributed data processing.

  • Experience with cloud data platforms (Databricks, BigQuery, Snowflake, or similar).

  • Solid understanding of data modeling patterns (dimensional modeling, medallion architecture).

  • Experience with event streaming systems (Pub/Sub, Kafka, or similar).

  • Familiarity with GCP or other major cloud platforms.

  • Track record of building reliable, scalable pipelines in production.

Bonus if you have:
  • Experience with Databricks Asset Bundles or similar deployment frameworks.

  • Background in ML data pipelines: feature engineering, embedding generation, model serving data.

  • Familiarity with Protocol Buffers or other schema evolution tools.

  • Experience with vector databases and embedding workflows.

  • Background in marketing data: attribution, customer analytics, campaign tracking.

  • Experience with e‑commerce data sources (Shopify, Klaviyo, marketing platforms).

Our Stack
  • Data Processing: Databricks, Apache Spark, PySpark, dbt

  • Event Streaming: GCP Pub/Sub, Protocol Buffers

  • Storage: BigQuery, AlloyDB (PostgreSQL), Cloud Storage

  • ML/AI Data: pgvector, embedding pipelines, LLM training data

  • Infrastructure: GCP, Terraform, Kubernetes, GitHub Actions

  • Languages: Python 3.11, SQL

Why Join Us?
  • Join an innovative, fast‑growing startup building cutting‑edge AI marketing solutions.

  • Make a meaningful impact by shaping the platform's user experience, design identity, and overall success.

  • Dynamic environment with opportunities for real ownership, learning, and growth.

  • Competitive salary and support for professional development.

This is a hybrid role, with 4 days per week in our downtown Toronto office.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Engineer
Staff Data Engineer

Loblaw Digital • Toronto

On-site
CAD 198,000 - 268,000
Data Engineer
Data Engineer

Financeit • Toronto

On-site
CAD 75,000 - 95,000
Hybrid workplace
Competitive salary with bonus
Comprehensive benefits
+3
Senior Software Engineer
Senior Software Engineer

Drive Capital • Toronto

On-site
CAD 100,000 - 130,000
Competitive salary
Support for professional development
Dynamic work environment
Jr Data Engineer
Jr Data Engineer

Citylitics • Toronto

Hybrid
CAD 70,000 - 90,000
Senior Data Engineer
Senior Data Engineer

Socket.dev • Toronto

Hybrid
CAD 110,000 - 160,000
Flexible work arrangements
Generous vacation policy
Customizable benefits
+1
Lead Data Engineer / Data Platform Lead
Lead Data Engineer / Data Platform Lead

Princeton IT Services, Inc • Toronto

On-site
CAD 120,000 - 180,000
Senior Integrations Developer / AI Data Integration Engineer
Senior Integrations Developer / AI Data Integration Engineer

Engineered Intelligence Inc. • Mississauga

Hybrid
CAD 90,000 - 120,000
Competitive compensation and benefits
Flexible hours
Autonomy & Growth opportunities
Senior Data Engineer - Machine Learning & Data Platforms - REMOTE
Senior Data Engineer - Machine Learning & Data Platforms - REMOTE

TEEMA Solutions Group • Toronto

Remote
CAD 100,000 - 130,000
Data Engineer – Platform
Data Engineer – Platform

United States Digital Space LLC • Toronto

Hybrid
CAD 103,000 - 148,000
Extended health and dental coverage
Retirement savings plans
Monthly meal allowance
+4
Data Engineer Toronto
Data Engineer Toronto

Konrad • Toronto

On-site
CAD 90,000 - 125,000
Retirement Planning
Parental Leave Program
Flexible Working Hours
+4