Senior Data Engineer, Blockchain data and/or NLP pipelines

Unchain Data

Northern (KY)

Hybrid

USD 130,000 - 190,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inca Digital, a veteran-owned data and intelligence company, is seeking experts to design and operate end-to-end ETL/ELT pipelines and NLP-LLM processing over unstructured text. You will build production-grade APIs, dashboards, and automated alerts by sourcing data from databases, APIs, files, blockchain nodes, and social feeds.

Join a remote-first team collaborating with R&D, data scientists, and investigations to translate intelligence requirements into scalable client-facing products,

Qualifications

  • Advanced Python and SQL.
  • Production ETL/ELT pipelines with orchestrators (Airflow, Dagster, Prefect, Step Functions) including backfills, idempotency, retries, schema drift, and on-call.
  • Hands-on with PostgreSQL/RDS and object storage (S3).
  • Serving data to consumers: APIs (Kong OSS, FastAPI) and dashboards.
  • Docker, Git, CI/CD (GitHub Actions, AWS CodeBuild/CodePipeline), and IaC (Terraform, CDK).

Responsibilities

  • Design, build, and operate end-to-end batch and incremental ETL/ELT pipelines, turning raw data into production-grade APIs, dashboards, and automated alerts.
  • High-throughput data acquisition from diverse sources: relational databases, third-party and vendor APIs, flat files, blockchain nodes and indexers, and social and web text at scale.
  • NLP and LLM-assisted processing over unstructured text — entity and claim extraction, classification, deduplication, and enrichment feeding downstream risk models.
  • Model and query data across relational, object, and graph (Neo4j) stores, choosing the right one for the access pattern rather than defaulting to a favorite.
  • Provenance and reproducibility: raw captures immutable, datasets versioned, and any published finding reproducible as of the date it was made.
  • Data quality as a first-class deliverable: validation, lineage, reconciliation, and freshness monitoring.
  • Partner with the Head of R&D, engineers, data scientists, and our investigations team to translate intelligence requirements into scalable client-facing products.
  • Contribute to technical design and architecture in a written, argued design process.

Skills

Python
SQL
Airflow
Dagster
Prefect
PostgreSQL
API design
Docker
Git
CI/CD
Terraform
Kubernetes
FastAPI
S3

Tools

Kong OSS
Neo4j
S3

Job description

About Us

Inca Digital is a veteran-owned data and intelligence company specializing in digital-asset analytics for exchanges, financial institutions, regulators, and blockchain ecosystems. Our technology and expertise provide clarity across crypto markets, tracking blockchain transactions, liquidity movements, and illicit finance, helping clients identify risk, enhance transparency, and improve decision-making. Inca's infrastructure fuses structured and unstructured data from blockchains, exchanges, social networks, and financial markets. The result is a powerful analytics engine that supports ecosystem monitoring, market surveillance, and counter-illicit finance intelligence across digital-asset networks. Inca operates as a fast-paced, nimble, global, and remote technology company. We leverage an asynchronous-first workflow, try to minimize time spent on meetings, and believe in open debate and logic over authority. Work alongside some of the sharpest minds in the world, including intelligence analysts, defense veterans, data engineers, quant researchers, linguists, and more.

Domain Expertise
  • Blockchain Data: Direct, hands-on experience treating ledger data as standard data structures. You have pulled from RPC endpoints, archive nodes, or indexers; decoded logs and ABIs; managed chain reorgs and protocol edge cases; and applied UTXO vs. account-based models in production environments.
  • NLP & LLMs: Proven track record shipping production-grade text processing workflows—including extraction, classification, embeddings, vector search, or LLM-in-the-loop enrichment. You possess a clear, practical understanding of model evaluation, cost optimization, and failure mode mitigation.

We expect deep expertise in one of these domains and working proficiency in the other. Please highlight your primary focus area in your application.

What You'll Own
  • Design, build, and operate end-to-end batch and incremental ETL/ELT pipelines, turning raw data into production-grade APIs, dashboards, and automated alerts.
  • High-throughput data acquisition from diverse sources: relational databases, third-party and vendor APIs, flat files, blockchain nodes and indexers, and social and web text at scale.
  • NLP and LLM-assisted processing over unstructured text — entity and claim extraction, classification, deduplication, and enrichment feeding downstream risk models.
  • Model and query data across relational, object, and graph (Neo4j) stores, choosing the right one for the access pattern rather than defaulting to a favorite.
  • Provenance and reproducibility: raw captures immutable, datasets versioned, and any published finding reproducible as of the date it was made.
  • Data quality as a first-class deliverable: validation, lineage, reconciliation, and freshness monitoring.
  • Partner with the Head of R&D, engineers, data scientists, and our investigations team to translate intelligence requirements into scalable client-facing products.
  • Contribute to technical design and architecture in a written, argued design process.
Requirements
  • Advanced Python and SQL.
  • Proven experience building and operating production ETL/ELT pipelines under an orchestrator (Airflow, Dagster, Prefect, Step Functions, or equivalent), including backfills, idempotency, retries, schema drift, and on-call for your own pipelines.
  • Hands-on with PostgreSQL/RDS and object storage (S3).
  • Serving data to consumers: building and versioning APIs (Kong OSS, FastAPI, or similar) and/or feeding dashboards.
  • Docker, Git, CI/CD (GitHub Actions, AWS CodeBuild/CodePipeline), and IaC (Terraform, CDK, or equivalent).
Nice to Have
  • Streaming and near-real-time architectures (Kafka, AWS SQS/SNS).
  • Document stores (MongoDB) and analytical engines or warehouses (ClickHouse, DuckDB, Snowflake, Athena).
  • Graph data modeling and querying (Neo4j/Cypher or comparable).
  • Transformation and data-quality tooling (dbt, Great Expectations, Soda, OpenLineage).
  • Vector stores (pgvector, Qdrant, OpenSearch).
  • Go; Kubernetes.
  • Domain expertise in:
    • Blockchain ledger data, smart contracts, or web3 languages (Solidity, Vyper, Go, Bitcoin Script)
    • NLP techniques, LLM integrations, or model feature engineering
    • Financial services, trading venues, or regulatory frameworks (SEC, CFTC, FinCEN)
    • OSINT, dark web, or social platform data collection
Why Join Us
  • We operate in a fully remote, high-trust environment and prioritize impact and delivered results over hours logged.
  • We offer competitive compensation and healthcare stipend.
  • We invest in our people through childcare stipends and dedicated employee assistance resources.
  • Join a thought leader at the intersection of national security and digital assets, working with cutting-edge tech that defines the field.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer, Blockchain data and/or NLP pipelines
Senior Data Engineer, Blockchain data and/or NLP pipelines

Inca Digital • Northern (KY)

Hybrid
USD 120,000 - 180,000
Healthcare stipend
Childcare stipend
Employee assistance resources
Software Engineer, Data Infrastructure & Graph Systems
Software Engineer, Data Infrastructure & Graph Systems

Unchain Data • Northern (KY)

Hybrid
USD 140,000 - 200,000
Software Engineer, Data Infrastructure & Graph Systems
Software Engineer, Data Infrastructure & Graph Systems

Inca Digital • Northern (KY)

Hybrid
USD 120,000 - 180,000
Fully remote
Healthcare stipend
Childcare stipends
+1
Software Engineer, Data Infrastructure & Graph Systems
Software Engineer, Data Infrastructure & Graph Systems

Inca Digital, Inc. • United States

On-site
USD 140,000 - 210,000
Remote work
Healthcare stipend
Childcare stipend
+1
Software Engineer, Data Infrastructure & Graph Systems
Software Engineer, Data Infrastructure & Graph Systems

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
Remote work
Healthcare stipend
Employee assistance resources
+1
Senior Data Engineer, Blockchain data and/or NLP pipelines
Senior Data Engineer, Blockchain data and/or NLP pipelines

Far Coder • Northern (KY)

Hybrid
USD 25,000 - 40,000
Senior Data Engineer: Remote Blockchain & NLP Pipelines
Senior Data Engineer: Remote Blockchain & NLP Pipelines

Inca Digital • Northern (KY)

Hybrid
USD 120,000 - 180,000
Healthcare stipend
Childcare stipend
Employee assistance resources
Software Engineer, Data Infrastructure & Graph Systems
Software Engineer, Data Infrastructure & Graph Systems

Far Coder • Northern (KY)

Hybrid
USD 25,000 - 40,000
Senior Data Engineer, Data Cloud
Senior Data Engineer, Data Cloud

chainalysis-careers • New York (NY)

Hybrid
USD 130,000 - 160,000
R&D Data Engineering Intern
R&D Data Engineering Intern

Inca Digital Federal LLC • United States

On-site
USD 60,933 - 72,986