Coragrid is actively hiring a Software Engineer - Data Infrastructure to build the production pipelines, data quality systems, and lineage tools behind private-company intelligence.
Coragrid builds data infrastructure for private markets. Our goal is to bring public-market level data fidelity to private companies: official financials, registry filings, web evidence, source-linked outputs, and retrieval-ready datasets.
We care about speed, truthfulness, and low-friction tools. The product is built for investment teams, but the engineering culture is terminal-first: clean CLIs, direct workflows, Linux-native development, and no unnecessary layers.
This role sits at the intersection of software engineering, data infrastructure, and product-quality data workflows. You will help build the systems that acquire raw public and commercial data, transform it into reliable company intelligence, evaluate quality, preserve lineage, and deliver it into APIs, retrieval workflows, exports, and internal tools.
The ideal candidate combines strong software fundamentals with curiosity about data systems. You do not need to be a senior distributed systems engineer already, but you should care about clean code, repeatable pipelines, careful debugging, and building systems that make data quality visible.
- Production data pipelines: Build workflows that acquire, parse, normalize, enrich, evaluate, and deliver private-company data at scale.
- Data lineage and visibility: Create tools that show where data came from, how it changed, what quality checks ran, and where it is used.
- Source-linked datasets: Preserve evidence, URLs, documents, timestamps, raw source context, and quality signals next to structured outputs.
- Internal data refinement tools: Build focused systems that make huge amounts of information easy to store, access, inspect, and reuse.
- Quality evaluation: Help design checks, monitoring, and feedback loops that improve accuracy, freshness, completeness, and reliability.
- Platform reliability: Improve the systems around PostgreSQL, Qdrant, object storage, queues, containers, and deployment workflows.
- 1+ years of professional software engineering, data engineering, data science, or strong equivalent project experience.
- Programming fundamentals: Ability to write clear, tested code in Rust, Go, Python, TypeScript, Java, C#, Scala, or a similar language.
- Data systems basics: Experience working with structured data, SQL, data modeling, ETL/ELT workflows, or production data pipelines.
- Database experience: Experience with PostgreSQL or comparable relational databases such as MySQL and MariaDB.
- Engineering workflow: Good understanding of Git, testing, code review, CI, build, deployment, logs, and monitoring.
- Hands-on experience with Rust, Python, or backend services built with frameworks such as Axum, FastAPI, or similar.
- Experience building Rust-native CLI/TUI tools, shell workflows, or automation for data review, operations, and data-quality checks.
- Experience with Docker, Kubernetes, AWS, or containerized deployments.
- Experience with retrieval systems, vector databases, data lakes, object storage, or search/indexing infrastructure.
- Experience with observability and data exploration tools such as Grafana, Superset, notebooks, or comparable systems.
- Basic understanding of statistics, machine learning workflows, or evaluation methods for noisy real-world data.
- Strong fundamentals: You can turn ambiguous data problems into clear code, tests, and production behavior.
- Data quality mindset: You care about traceability, source evidence, idempotent writes, and clear failure modes.
- Backend curiosity: You want to move between services, APIs, database queries, ingestion jobs, infra, and integration work.
- Linux and CLI fluency: You prefer direct tools, shell workflows, logs, tests, and reproducible commands over heavy dashboards and bloat.
- Communication: You can concisely explain what broke, what changed, what evidence you found, and what tradeoffs remain.
- Systems taste: You want to build software that is fast, simple, observable, and precise for both humans and machines.
- Infra: Linux, Docker, Kubernetes, AWS, and GitHub Actions.
- Interfaces: API, CLI, TUI, exports, retrieval, pipelines, internal tools.
- Terminal-first: We optimize for fast local feedback, clear commands, readable logs, and systems that can be operated without ceremony.
- Zero bloat: We avoid unnecessary abstractions, heavy processes, and vague enterprise software habits.
- Source-traceable: Data products must be inspectable. If a claim exists, the source should be close by.
- Small sharp releases: We prefer tight diffs, tests, reviewable changes, and production behavior we can explain.
- Compensation based on skillset, role scope, and experience.
- High ownership: You will shape important product and data-infrastructure decisions early, not just close tickets.
- Hard technical problems: Dealing with billions of data points in private-market workflows are messy in exactly the right way.
- Modern tools: You get strong local hardware, serious developer tooling, and a team that values speed without sacrificing correctness.
- Founder-level visibility: You will work close to product direction, customer needs, and the technical architecture from day one.