Mach aus dieser Rolle ein Vorstellungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.
Aether Biomedical is seeking a remote data platform engineer to design, build, and maintain an AI-ready data platform. You will ingest, process, and enrich publisher content, enabling reliable access for multiple consumers.
Key responsibilities include building multi-stage pipelines, using S3 as storage, and implementing vector search with Weaviate. Proficiency in Python, FastAPI, and PostgreSQL is essential; remote role with onboarding in Warsaw.
Location: remote with the first day onboarding in Warsaw and occasional visits once per quarterRate: 170 pln/h on b2b
We are building CaaS (Content as a Service) a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data.The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.
The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.
Data Engineering & PipelinesBuild and maintain multi-stage data ingestion pipelinesDesign and implement idempotent, restartable batch processing workflowsUse S3 as core storage layer for raw and processed dataImplement pipeline stages including:Content ingestion and book identity assignmentPDF-to-markdown conversion (AI OCR)Table of contents and structure extractionHierarchical chunkingEmbedding generationAI / LLM ProcessingUse LLMs and OCR models to extract structured data from PDFsDesign prompts and context strategies for consistent outputsGenerate structured metadata and enrich content for downstream use casesData Storage & ConsistencyMaintain PostgreSQL (Aurora) as system of recordDesign and maintain SQL schemas and versioned migrationsEnsure data consistency across:S3PostgreSQL (Aurora)Vector database (Weaviate)Implement reconciliation logic across distributed systemsRetrieval & Vector SearchWork with Weaviate for vector search and semantic retrievalSupport RAG-based applicationsDesign data organization strategies (by subject, country, and client)APIs & IntegrationBuild REST APIs using FastAPIExpose content as a service for multiple downstream applicationsIntegrate with internal and external systemsEngineering PracticesWrite strongly typed Python code (mypy)Follow CI/CD processes with automated checks (ruff, pytest)Work across dev / staging / production environmentsDebug distributed data inconsistencies