Software Engineer, Data Mining

NationGraph Inc.

Toronto

On-site

CAD 90,000 - 140,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NationGraph is building a data and intelligence layer for the public sector. We are seeking a Software Engineer, Data Mining to own the systems that acquire public-sector information from across the internet at massive scale.

You will build and operate hundreds of thousands of scrapers, designing scalable abstractions and ensuring data quality while leveraging AI and agents to improve discovery and extraction. Toronto-based, fully remote options are not specified.

Qualifications

  • Experience building production web crawlers, scraping systems, browser automation, or large-scale external data pipelines.
  • Strong coding skills in Python, Go, and TypeScript for backend/systems work.
  • Understanding of scraping modern websites, including JS rendering, rate limits, proxies, and authentication.
  • Experience with distributed systems, orchestration, queues, retries, and observability.

Responsibilities

  • Own our scraping infrastructure end-to-end: create, deploy, schedule, monitor, and maintain hundreds of thousands of scrapers.
  • Build abstractions to scale toward millions of sources with minimal engineering effort.
  • Work across web crawling, browser automation, APIs, PDFs, and legacy data sources.
  • Enable AI-driven scraping with LLMs and agents to discover sources and repair failures.
  • Scale systems for orchestration, concurrency, retries, backfills, and cost management.
  • Expand coverage globally across the U.S., Canada, and beyond.

Skills

Python
Go
TypeScript
Distributed systems
Web crawling

Tools

PostgreSQL
Redis
Docker
Kubernetes

Job description

Software Engineer, Data Mining
About NationGraph

NationGraph is building the data and intelligence layer for the public sector.

  • More than 110,000 state and local government agencies across the U.S. independently publish information about:

    • How they operate

    • What they buy

    • Who they work with

    • What problems they are trying to solve

  • That information is fragmented across millions of websites, documents, databases, procurement systems, meeting records, and public records.

  • NationGraph turns that information into structured, connected, actionable intelligence for businesses selling to government.

  • Founded in 2024, NationGraph is dedicated to making uncommon knowledge common, because public data should actually be public.

The Role

We’re looking for a Software Engineer, Data Mining to own one of the most important technical problems at NationGraph: building the systems that acquire public-sector information from across the internet at massive scale.

Our goal is to operate hundreds of thousands, and eventually millions, of scrapers covering every level of government across the U.S. and Canada, and eventually worldwide.

This is not a role focused on manually building individual scrapers. You’ll own the infrastructure, abstractions, and automation that allow us to create, deploy, monitor, and maintain an enormous fleet of scrapers reliably.

You’ll work across:

  • Web crawling and scraping

  • Browser automation

  • Distributed systems

  • Data extraction

  • Infrastructure and orchestration

  • LLMs and agents

  • Monitoring and observability

What You’ll Do
  • Own our scraping infrastructure end-to-end

    • Build systems for creating, deploying, scheduling, monitoring, and maintaining hundreds of thousands of scrapers.

    • Design abstractions that allow us to scale toward millions of sources without scaling engineering effort linearly.

  • Build for the messy internet

    • Work across government websites, APIs, procurement systems, PDFs, spreadsheets, meeting records, and legacy systems.

    • Handle changing websites, undocumented APIs, rate limits, broken sources, and countless edge cases.

  • Make scraping a distributed systems problem

    • Build for orchestration, concurrency, retries, backfills, change detection, observability, cost management, and failure recovery.

    • Ensure we know when sources break, data disappears, or extraction silently becomes incorrect.

  • Use AI to rethink scraping

    • Work with our ML Research team to use LLMs and agents to:

      • Discover new sources

      • Understand unfamiliar websites

      • Generate scraping logic

      • Detect source changes

      • Diagnose and repair failures

      • Validate extracted data

  • Build systems that get better with scale

    • Identify common platforms and patterns that can unlock thousands of government agencies at once.

    • Make new sources increasingly cheap and automated to onboard.

  • Expand our coverage globally

    • Help comprehensively map public-sector information across the U.S. and Canada.

    • Build the foundation to eventually acquire public-sector information worldwide.

You Might Be a Good Fit If
  • You’re an unusually strong engineer who enjoys figuring out how things work.

  • You’ve built production web crawlers, scraping systems, browser automation, or large-scale external data pipelines.

  • You’re strong in Python, Go, TypeScript, or another backend/systems language.

  • You understand the realities of scraping modern websites, including:

    • JavaScript rendering

    • Sessions and cookies

    • Rate limits

    • Proxies

    • Authentication

    • Changing schemas and websites

  • You understand distributed systems, including:

    • Orchestration

    • Queues and concurrency

    • Idempotency

    • Retries

    • Backfills

    • Observability

    • Failure recovery

  • You care deeply about data quality, correctness, and reliability.

  • You’re excited about using LLMs and agents to automate traditionally manual scraping work.

  • You naturally think about leverage: not how to scrape one website, but how to build a system capable of scraping the next 10,000.

  • You thrive in ambiguity and would rather build the system than be handed one.

We’re particularly interested in backgrounds spanning:

  • Alternative data

  • Quantitative research infrastructure

  • Search and crawling

  • AI data infrastructure

  • Knowledge graphs

  • Large-scale document processing

  • Data aggregation

None of these are requirements.

Our Engineering Stack
  • Backend: Python, Go, PostgreSQL

  • Infrastructure: Redis, Docker, Kubernetes

  • Frontend: React, TypeScript

  • AI / ML: LLMs, agents

Our stack will evolve, you’ll help decide how.

Why NationGraph
  • Own a foundational problem

    • A large part of this architecture still needs to be invented.

    • You’ll have significant ownership over how NationGraph discovers, acquires, represents, and serves public-sector information.

  • Work on a genuinely hard data problem

    • There is no single API for American government.

    • There are tens of thousands of institutions, millions of sources, inconsistent schemas, and enormous amounts of information buried in systems never designed for machines.

  • Build a real data moat

    • We believe a major long-term advantage in applied AI will come from proprietary context and data.

    • Government contains enormous amounts of valuable information that is technically public but practically inaccessible.

    • Your job is to change that.

  • Work with exceptional people

    • You’ll work closely with the CEO, CTO, and a small engineering and research team.

    • The team has backgrounds spanning high-scale infrastructure, quantitative finance, AI, and startups.

  • Have real ownership

    • We move quickly.

    • We operate with very little bureaucracy.

    • Engineers have significant ownership over technical decisions and product outcomes.

If the idea of building the data infrastructure to map and understand how government works sounds exciting, we’d love to talk.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer (Growth)
Software Engineer (Growth)

NationGraph Inc. • Toronto

On-site
CAD 90,000 - 150,000
Growth Engineer
Growth Engineer

NationGraph Inc. • Toronto

On-site
CAD 110,000 - 160,000
Software Engineer (Full Stack) - 6 Month Contract
Software Engineer (Full Stack) - 6 Month Contract

Citylitics • Toronto

On-site
CAD 100,000 - 140,000
Engineering Manager - AI Platform
Engineering Manager - AI Platform

NEOGOV • Toronto

Hybrid
CAD 199,000 - 284,000
Generous PTO
Remote working opportunities
401K Matching
+1
Software Engineer (Full Stack) - 6 Month Contract
Software Engineer (Full Stack) - 6 Month Contract

Citylitics Inc. • Toronto

Hybrid
CAD 85,000 - 120,000
Impactful public infrastructure work
Access to Generative AI tools
HybridToronto office
AI Infra Engineer
AI Infra Engineer

Urban Ridge Supplies • Toronto

Remote
CAD 120,000 - 180,000
Competitive Salary
Health Insurance (US Only)
Remote Work Environment
+2
Sr. Software Engineer, GenAI
Sr. Software Engineer, GenAI

The Consensus • Canada

On-site
CAD 246,000 - 309,000
Health, dental, vision insurance
Generous PTO
401(k)
+2
Solutions Architect, AI Systems
Solutions Architect, AI Systems

Viral-Natio • Toronto

On-site
CAD 130,000 - 150,000
Head of Data Infrastructure & Analytics (Fully Remote)
Head of Data Infrastructure & Analytics (Fully Remote)

Puffy • Quebec

On-site
CAD 180,000 - 220,000
Monthly performance bonuses up to 25%
Comprehensive medical, dental, and vision insurance
Generous Paid Time Off (PTO)
+2
Senior Analytics Engineer
Senior Analytics Engineer

Passage • Toronto

On-site
CAD 100,000 - 150,000