data engineer for unstructured data and AI pipelines

HireHi

United States

Remote

USD 60,000 - 90,000

Part time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

HireHi ищет опытного инженера по данным для создания приватного контекстного слоя над данными европейской группы, с возможностями обработки неструктурированных источников и интеграцией API. Роль предполагает работу над инжекторингом данных, пайплайнами, обеспечения аудита и безопасности, а также совместную работу в небольшой гибридной команде на part-time.

Кандидат должен владеть Python, SQL, Airflow и опытом работы с крупными API, приватности и GDPR, с упором на качественную Retrieval-augmented

Qualifications

  • 4+ года в data engineering, работа с неструктурированными данными.
  • Опыт интеграции нескольких API и исторической подкачки данных.
  • Работа с чувствительными данными в регулируемой среде.
  • Умение работать как единственный data engineer в маленькой команде.

Responsibilities

  • Настройка ingestion через call, email, Slack и мессенджеры с атрибуцией источника и откатом.
  • Backfill исторических данных (email, Slack, протоколы, презентации и портфели).
  • Разработка парсинга документов (PDF, сканы, таблицы, презентации).
  • Реализация идентификации и выравнивания сущностей между инструментами и CRM.
  • Построение конвейеров chunking/embeddings и загрузка в графовую/векторную базу.
  • Обеспечение инкрементной синхронизации без повторной crawl-обработки.
  • Добавление области доступа и происхождения к каждой записи.
  • Осуществление PII-декларирования, редактирования и retention логики.
  • Оркестрация пайплайнов через Airflow/Step Functions и мониторинг источников.
  • Контроль расходов и задержек через батчинг и хранение.
  • Подготовка runbooks для handover.

Skills

Python
SQL
Airflow
APIs
Data ingestion

Tools

pgvector
OpenSearch
Pinecone
AWS
GCP

Job description

Описание:

The project builds a private, access-scoped context layer over a European private investment group's data, with AI skills and agents built on top.

Задачи:
  • Set up call, email, Slack, and messenger ingestion with speaker attribution and reversible opt-out
  • Backfill historical email, Slack, board protocols, decks, and portfolio updates, ensuring they are parsed, deduplicated, and correctly dated
  • Build document parsing for PDFs, scanned board packs, spreadsheets, slide decks, and forwarded attachments
  • Implement identity and entity resolution across communication tools, calendars, portfolio companies, and CRM records
  • Build chunking and embedding pipelines and load vector and graph stores behind the architect-defined ontology
  • Implement incremental sync through the connector layer, handling edits and deletions without full re-crawls or silent drift
  • Attach access scope and provenance to every record during ingestion to support downstream permission-aware retrieval and audits
  • Run PII detection, redaction, and retention logic; provide the client's security function with evidence of what is stored, where, and for how long
  • Orchestrate monitored, repeatable pipelines with Airflow, Step Functions, or equivalent, and alert when a source stops flowing
  • Control cost and latency at volume through batching, incremental embedding, and storage tiering; report unit economics
  • Write runbooks so the client's team can operate the system after handover
Требования:
  • 4+ Years in data engineering, including unstructured or semi-structured data work beyond warehouse modelling
  • Demonstrated experience integrating multiple third-party APIs into a coherent store, including historical backfill
  • Experience handling sensitive personal data in a regulated or security-sensitive environment
  • Comfortable working as the only data engineer in a small 2.5-FTE pod, at part-time allocation, without hand-holding
  • Strong Python and solid SQL
  • Experience with unstructured-data pipelines for transcripts, mail, chat, and documents, including parsing, normalisation, and deduplication
  • Experience with chunking strategies, vector stores such as pgvector, OpenSearch, or Pinecone-class systems, and graph-store loading
  • Experience integrating APIs and connectors at scale, including Google Workspace or M365, Slack, and CRM; rate limits, pagination, incremental cursors, and webhooks
  • Experience with entity resolution or record linkage, deterministic and fuzzy, without a clean shared key
  • Experience with Airflow, Step Functions, or equivalent orchestration and idempotent, restartable jobs
  • AWS and/or GCP data stack experience and comfort with private or VPC deployments
  • Practical knowledge of PII detection, redaction, encryption, and retention
  • Clear written English; able to prepare handover documents and work asynchronously in a small distributed pod
  • Knowledge of GDPR as applied to employee-generated data and EU data residency across multiple jurisdictions
  • Knowledge of data lineage, provenance, and audit patterns
  • Understanding of how retrieval quality depends on ingestion quality and sufficient RAG knowledge to make upstream choices
  • Будет плюсом: experience building pipelines feeding an LLM or retrieval system, Well-Architected security and cost practices, awareness of financial-services expectations
Условия:
  • Part-time engagement, 20 to 30 hours a week; allocation may flex above 0.5 FTE during Capture and Connect and settle back afterwards
  • Approximately eight to ten two-week sprints overall, with workload front-weighted to the first four or five Active project; start ASAP
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer (UA/RU Language speaking)
Data Engineer (UA/RU Language speaking)

Neurons Lab • Spain (TX)

On-site
USD 60,000 - 120,000
ai engineer for internal AI tools
ai engineer for internal AI tools

HireHi • United States

Hybrid
USD 101,000 - 169,000
Learning budget
Mental health sessions
On-site workshops
+1
.net developer for regulated financial firms
.net developer for regulated financial firms

HireHi • United States

Remote
USD 120,000 - 180,000
data engineer data integration
data engineer data integration

HireHi • United States

Remote
USD 120,000 - 180,000
full stack developer for data-centric products
full stack developer for data-centric products

HireHi • United States

Remote
USD 140,000 - 190,000
data engineer in fintech
data engineer in fintech

Enfint • United States

On-site
USD 120,000 - 180,000
Provident Fund
Annual learning budget
€150 Monthly Wolt allowance
+4
data engineer for retail platforms
data engineer for retail platforms

HireHi • United States

On-site
USD 180,000 - 240,000
data engineer in regulated financial services
data engineer in regulated financial services

HireHi • United States

Remote
USD 120,000 - 190,000
Part-Time Data Engineer: Unstructured Data & AI Pipelines
Part-Time Data Engineer: Unstructured Data & AI Pipelines

HireHi • United States

Remote
USD 60,000 - 90,000
DevOps Engineer Data
DevOps Engineer Data

HireHi • United States

On-site
USD 140,000 - 190,000