Data Engineer

CLR3

Toronto

Hybrid

CAD 100,000 - 160,000

Full time

41 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

CLR3 is hiring a data engineering role to build and run large-scale decoding pipelines for on-chain history, producing reliable Parquet data with strong lineage and reproducibility. You will model schemas across protocols and ensure data correctness through validations and checksums.

You will own end-to-end pipelines, including deployment, backfills, and handling production issues, while collaborating with teams handling schema coverage and catalog updates.

Qualifications

  • Production data pipelines at scale.
  • Strong SQL and at least one of Python, Rust or Go.
  • Familiarity with columnar formats such as Parquet.
  • Experience with DuckDB, Polars or Spark.
  • Care for data correctness: testing and validation.
  • Comfort owning pipelines in production and on-call responsibility.
  • Able to work from Toronto office part of the week.

Responsibilities

  • Design and run large-scale decoding and backfill pipelines.
  • Model typed schemas for instructions, events and state across dozens of protocols.
  • Build validation that catches bad data before customers do.
  • Keep versioning, checksums, manifests and lineage docs accurate on every delivery.
  • Tune storage layout and partitioning so files query fast in DuckDB and Polars.
  • Add new protocols and chains to the catalogue.
  • Support customer questions about schemas and coverage

Skills

SQL
Python
Rust
Go
Pipelines
Parquet
DuckDB
Polars
Spark
Testing

Tools

Orchestration tools

Job description

Build the pipelines behind datastore: decoding years of on-chain history into clean, versioned Parquet that researchers can trust.

datastore sells something unusual: files, not API access. Customers download decoded on-chain history as Parquet and run their own queries. That only works if the data is actually right, which makes correctness, lineage and reproducibility the product.

You will build and run the pipelines that decode Solana and Hyperliquid history at scale: backfills over billions of rows, schema design, checksums, manifests and the quality checks that let a quant trust a file they did not produce themselves.

What you will do
  • Design and run large-scale decoding and backfill pipelines
  • Model typed schemas for instructions, events and state across dozens of protocols
  • Build validation that catches bad data before customers do
  • Keep versioning, checksums, manifests and lineage docs accurate on every delivery
  • Tune storage layout and partitioning so files query fast in DuckDB and Polars
  • Add new protocols and chains to the catalogue
  • Support customer questions about schemas and coverage
What we are looking for
  • Experience building production data pipelines at meaningful scale
  • Strong SQL plus one of Python, Rust or Go
  • Real familiarity with columnar formats, ideally Parquet, and query engines like DuckDB, Polars or Spark
  • Care for data correctness: testing, validation and reconciliation
  • Comfort owning pipelines in production, including when they break at night
  • Able to work from our Toronto office part of the week
Nice to have
  • Experience with blockchain data or other messy, high-volume event streams
  • Familiarity with warehouse ecosystems your customers use, like Snowflake, BigQuery or ClickHouse
  • Background in quantitative research support or backtesting infrastructure
  • Experience with orchestration tools and with knowing when a cron job is enough
Who you are
  • You think an unverified number is worse than no number
  • You write pipelines you would be happy to debug at 2am, so they rarely need it
  • You like schemas that make the next person's query obvious
  • You get satisfaction from a backfill that reconciles to the last row
How we hire
  • 1 Intro call with an engineer, about 30 minutes
  • 2 Short take-home assignment working with a real decoded dataset
  • 3 Technical conversation about your assignment and pipelines you have run
  • 4 Conversation with the founders
  • 5 Offer
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Rust Engineer, Trading Infrastructure
Rust Engineer, Trading Infrastructure

CLR3 • Toronto

Hybrid
CAD 120,000 - 180,000
Software Engineer II
Software Engineer II

United States Digital Space LLC • Toronto

On-site
CAD 90,000 - 150,000
Senior Analytics Engineer
Senior Analytics Engineer

Passage • Toronto

On-site
CAD 100,000 - 150,000
Data Platform Engineer
Data Platform Engineer

Chair.com.pk • Toronto

On-site
CAD 90,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Aeolus Data • Canada

Hybrid
CAD 95,000 - 125,000
Software Engineer
Software Engineer

CLR3 • Toronto

Hybrid
CAD 110,000 - 170,000
Data Engineer
Data Engineer

Loopio • Canada

Remote
CAD 95,000 - 135,000
Health benefits
Parental leave
Stock options
+5
S Staff Software Engineer, Product Risk Stripe via Greenhouse Toronto 8122 data foundations View role
S Staff Software Engineer, Product Risk Stripe via Greenhouse Toronto 8122 data foundations View role

Nubeero Limited • Toronto

On-site
CAD 140,000 - 180,000
Data Engineering Developer
Data Engineering Developer

Osedea • Montreal (administrative region)

Hybrid
CAD 90,000 - 120,000
Competitive salary RRSP
Flexible hours
8 weeks remote work
+5
Senior Data Engineer, NimbleRx
Senior Data Engineer, NimbleRx

Swoop Airlines and Aviation • Toronto

On-site
CAD 120,000 - 150,000