A complete application in a minute — tailored resume and cover letter, ready to send.
Algocor is seeking a Senior Data Engineer to own high-volume data ingestion and time-series infrastructure behind our quant and AI systems. You will design, build, and operate the data layer connecting market data providers, exchange APIs, and cloud services.
You will implement robust ETL/ELT pipelines, streaming data from WebSocket and API sources, and production-grade time-series databases such as QuestDB, kdb+, ClickHouse, or TimescaleDB. Strong Python and SQL skills are essential.
Algocor is looking for aSenior Data Engineerto own the high-volume data ingestion and time-series data infrastructure behind our quant and AI systems.
We are building a trading system where research, execution, market data, and an LLM-based agent layer operate on the same infrastructure. For this system to work reliably, the data layer needs to be more than a pipeline. It needs to be well-structured, observable, recoverable, and trusted by both quant systems and AI agents.
You will design, build, and operate the data ingestion, transformation, storage, and access layer that connects external market-data providers, exchange and broker APIs, on-premise and cloud-based systems.
Your work will include:
Building and maintaining high-volume ETL/ELT pipelines from market-data providers such as Pyth, Databento, exchange APIs, and broker APIs into our on-prem stack
Designing and operating streaming data pipelines from WebSocket and API sources
Owning production-grade time-series database design and operations using systems such as QuestDB, kdb+, ClickHouse, TimescaleDB, or similar
Designing data structures for tick data, OHLCV, symbols, derived signals, and internal datasets
Making decisions around partitioning, retention, compression, schema evolution, query performance, and storage strategy
Designing the boundary between on-prem systems and AWS-based cloud components
Deciding what gets calculated where, how data is synchronized, and how sync health is monitored
Handling streaming failure scenarios such as reconnect logic, replay, backfill, duplicates, out-of-order events, late-arriving data, and gap detection
Writing production-grade Python services and APIs, including FastAPI, to expose clean and validated data to internal systems and the AI layer
Owning validation rules, data quality checks, observability, alerting, and recovery procedures
Building agent-facing data access tools such as query interfaces, document retrieval flows, and dataset endpoints
Documenting datasets, schemas, access rules, operational assumptions, and failure modes so the rest of the team can build confidently on your work
This is a role where you will be expected to scope, build, ship, monitor, and improve the systems you own.
Our AI layer is only as reliable as the data infrastructure underneath it.
In this role, your work will directly shape how confidently we can use data across quant research, execution systems, internal tools, and AI agents.
You will be close to the architecture, the data, and the people building on top of it. This is a high-ownership role in a small, focused team where individual contribution is visible.
You will have:
End-to-end ownership of the data ingestion and storage layer
Direct collaboration with the Engineering and quant team on architecture
A modern stack with real engineering problems
The opportunity to build infrastructure that directly supports quant and AI systems
We are looking for a senior engineer who can make independent decisions around data ingestion, storage, streaming reliability, and production data infrastructure. You are likely to be a strong fit if you have:
You have built or operated data systems in production, not just experimental projects, dashboards, or offline analytics pipelines.
You understand what changes when data volume grows significantly - including throughput, batching, partitioning, storage cost, write performance, backfill strategy, and operational monitoring.
You have worked with continuously flowing data from APIs, WebSockets, message brokers, or event streams. You understand replay, gap detection, ordering, duplicates, late data, and recovery.
You have hands-on experience with time-series or high-volume analytical databases such as QuestDB, kdb+, ClickHouse, TimescaleDB, or similar systems. You understand data modeling, partitioning, retention, query performance, and operational trade-offs.
You write production-grade Python and strong SQL. You can build maintainable pipelines, services, and APIs that other systems depend on.
You are comfortable with Docker, Linux, Git, CI/CD, monitoring, alerting, incident response, documentation, and owning what happens after something is shipped.
You can take a research, business, or product need and turn it into a practical technical specification without overcomplicating the process.
Büdotek Teknopark, Istanbul