Principal Research Software Engineer - Bioinformatics
Bridge Analytics is seeking a Principal Research Software Engineer to build the applications and data that researchers actually touch inside the Bridge Analytics Environment (BAE) - a Trusted Research Environment on Google Cloud Platform where the world’s leading neuroscientists discover data, build cohorts, and generate publication-quality analyses. You will own the researcher-facing applications (starting with our no-code Visualizer) and the curated, harmonized datasets and pipelines beneath them - the full path from raw multi-modal data to the interactive tools that turn it into discovery for Parkinson's disease, autism, and bipolar disorder.
This is a senior, hands-on individual-contributor role. You build directly across the stack - from BigQuery pipelines to a Next.js frontend - and you coordinate contractor and consultant resources for data curation, pipeline implementation, and help desk, so you multiply your impact without managing a large team. We are looking for a data- and science-grounded engineer who can also ship a clean frontend: someone who makes datasets trustworthy and FAIR, and makes them usable through great software.
Key Responsibilities
- Researcher-Facing Applications: Own and extend the apps researchers use every day - the no-code Visualizer, cohort builder, and data catalog - across a FastAPI (Python) backend and a Next.js / TypeScript / Tailwind frontend, against live BigQuery and Dataplex.
- Agentic & Interactive Analytics: Evolve the agent layer (Google ADK + Gemini, with a model-agnostic path to other frontier models) that turns natural-language questions into interactive, publication-quality dashboards.
- Data Pipelines: Design and operate the ingestion, curation, and harmonization pipelines (Cloud Composer/Airflow, Dataform, BigQuery) that land multi-modal datasets - clinical, imaging, sensor, and 'omics - from external data-coordinating centers and partner APIs, with the LinkML/Pydantic schemas that define them.
- Curated, Trustworthy Data: Keep data curation moving and release-ready - coordinate contractor and consultant curation resources, set standards, and implement data quality, validation, and lineage so every released dataset is FAIR and discoverable.
- Scalable Scientific Compute: Get workflow engines (e.g., Nextflow/nf-core) running elastically on GCP (Batch/GKE) so researchers can run imaging and 'omics pipelines at real scale.
- Portfolio Scale & Integrations: Make the apps and data work across a growing multi-study portfolio - replace mocks with real integrations (Cloud Identity, Dataplex, VPC-SC) and build out end-to-end tests (Playwright) as the surface grows.
- Data Graph: Help model the data graph (Dataplex + BigQuery graph) so both researchers and agents can traverse well-described data, in partnership with the AI/Agentic and Infrastructure leads.
Who You Are
- You are a data- and science-grounded engineer at core - you've built production pipelines on real biomedical, clinical, or 'omics data - and you can also build and ship a clean React/Next.js frontend.
- You care that datasets are correct, documented, and FAIR, and you know how to make them usable through great software, not just queries.
- You thrive as a hands-on individual contributor - you'd rather build and set direction than manage a large team, and you know how to get leverage from contractor and consultant resources.
- You have startup DNA: broad ownership, fast shipping, and comfort coordinating vendors or contractors rather than waiting for a big team.
- You are eager to adopt and integrate modern AI tools into your daily workflow to multiply your impact - this is an AI-native team.
- You care about the mission and want your work in researchers' hands in weeks, not years.
Qualifications
Minimum Qualifications:
- Experience: 6+ years building production software and data systems, with the ownership to take a surface end-to-end.
- Data Engineering: Expert SQL and Python, with production experience on cloud data warehouses (BigQuery) plus orchestration (Airflow/Composer) and transformation frameworks (dbt/Dataform).
- Biomedical Data: Hands-on experience with biomedical, clinical, or 'omics data, and the scientific grounding to engage researchers as a peer.
- Startup / Early-Stage: A track record of thriving in high-ownership, early-stage environments - building broadly, shipping fast, and coordinating contractors or consultants.
Preferred Qualifications:
- Full-Stack / Frontend (strongly preferred): A Python backend (FastAPI/Flask) and a modern React framework (Next.js/TypeScript); building interactive data-visualization UIs.
- Scientific Compute: Workflow engines (Nextflow/nf-core, WDL, Snakemake) at scale on cloud; neuroimaging or 'omics data.
- AI / Agent Frameworks: LLM/agent frameworks (Google ADK, Claude Agent SDK, LangGraph); building agentic or data-visualization applications.
- Research Data Standards: FAIR principles and standards (FHIR, OMOP, GA4GH); schema modeling (LinkML, Pydantic); Trusted Research Environments.