C-BRAIN Data Engineer (Remote) - Neurology

Wustl

United States

Remote

USD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote work
Flexible schedule

Job summary

Washington University in St. Louis (WUSTL) seeks a C-BRAIN Data Engineer to design, build, and maintain the data infrastructure powering our AI biomedical tools.

You will own data ingestion, pipeline development, harmonization, and cloud operations to deliver analysis-ready data for research teams and AI experiments. You will work closely with the CTO, product managers, and data scientists to translate requirements into scalable pipelines using Spark, dbt, Airflow, and cloud services, while

Qualifications

  • Experience designing scalable data ingestion pipelines for diverse data sources.
  • Proficiency with ETL/ELT workflows using Spark, dbt, Airflow or equivalent.
  • Knowledge of data governance, access controls, and de-identification practices.
  • Familiarity with biomedical or neurodegeneration data is a plus.

Responsibilities

  • Designs, builds, tests, and maintains scalable data ingestion pipelines.
  • Develops and maintains ETL/ELT workflows with Spark, dbt, Airflow.
  • Implements monitoring, alerting, and error resolution for pipelines.
  • Collaborates with CTO and data scientists to translate requirements into pipelines.
  • Maintains version control and documentation for pipelines and infra.

Skills

ETL pipelines
Cloud platforms
Data pipeline development
Data governance
Version control

Tools

Apache Spark
dbt
Airflow

Job description

Location

Remote, US

Scheduled Hours

40

Position Summary

The C-BRAIN Data Engineer is a key technical member of the C-BRAIN team responsible for designing, building, and maintaining the data infrastructure that powers C-BRAIN's AI tools. Reporting to the C-BRAIN Chief Technology Officer (CTO), this role is responsible for all aspects of data ingestion, pipeline development, data harmonization, and cloud infrastructure management — ensuring that high-quality, analysis-ready data is available to C-BRAIN's AI tools and research teams. C-BRAIN is building an AI Biomedical Research Scientist platform that integrates diverse multi-institutional datasets (including NACC, ADNI, and consortium member data contributions). The Data Engineer will be central to building the technical infrastructure that makes this platform possible, working in close partnership with the CTO, the Senior Technical Product Manager, and external data science collaborators. This is not a standard data pipeline position. The Data Engineer is building the technical backbone of an AI biomedical research platform — infrastructure that must ingest and harmonize multi-modal neurodegeneration datasets at consortium scale and serve as the data foundation for agentic AI tools including InsightEngine and OpenScientist. The ideal candidate brings software engineering discipline, strong cloud platform experience, and demonstrated knowledge of neurodegeneration or biomedical research data. Domain knowledge is a prerequisite, not a nice-to-have; C-BRAIN-specific context will be provided, but neurodegeneration data experience and software engineering fundamentals will not.

Job Description

Primary Duties & Responsibilities:

Data Pipeline Development and Maintenance

  • Designs, builds, tests, and maintains scalable data ingestion pipelines to ingest consortium member datasets from diverse sources and formats into the C-BRAIN data infrastructure.
  • Develops and maintains ETL/ELT workflows using tools such as Apache Spark, dbt, Airflow, or equivalent; ensure pipelines are robust, well-documented, and auditable.
  • Implements automated pipeline monitoring and alerting; troubleshoot and resolves pipeline failures in a timely manner.
  • Works collaboratively with the CTO and data science teams to understand data requirements for AI tool development and translates those requirements into technical pipeline specifications.
  • Maintains version control for all pipeline code and infrastructure configurations; follows software engineering best practices including code review and documentation.
  • Integrates and processes multi-modal data including omics (genomics, transcriptomics, proteomics), neuroimaging (PET, MRI), longitudinal clinical records, and digital pathology — reconciling differences in data type, format, spatial resolution, and dimensionality into unified analytical frameworks.
  • Identifies where cross-modal integration produces genuine signal versus where it introduces noise or artifact; establishes ground truth benchmarks for downstream AI use.

Data Infrastructure and Cloud Operations

  • Manages and optimizes the C-BRAIN data infrastructure: storage accounts, computes resources, data lakes, and access controls.
  • Implements and maintains data access controls and permissions aligned with DUA requirements and WashU data governance policies.
  • Collaborates with the CTO on cloud architecture decisions; contributes to infrastructure planning for Phase 2 scale-up including foundation model compute requirements.
  • Monitors infrastructure costs, resource utilization, and performance; identifies and implements optimization opportunities.
  • Supports the deployment of C-BRAIN AI tools on cloud-based platforms; coordinates with technical teams on infrastructure requirements.
  • Ensures all data handling complies with DUA terms and applicable PHI de-identification requirements; implements, documents, and maintains de-identification workflows for each incoming dataset.
  • Uploads curated datasets to ADDI/AD Workbench and other designated repositories (NIAGADS, GP2, or equivalent) as directed; manages access controls within the platform to ensure data is accessible only by authorized users and tools.

Data Harmonization and Quality

  • Develops and implements data harmonization procedures to integrate datasets from multiple sources (
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Biomedical AI Data Platform Engineer
Remote Biomedical AI Data Platform Engineer

UNKNOWN • Missouri

Hybrid
USD 75,000 - 129,000
Remote Data Engineer for AI Biomedical Platform
Remote Data Engineer for AI Biomedical Platform

Washington University • St. Louis (MO), Northern (KY)

Hybrid
USD 120,000 - 180,000
Remote Data Engineer for AI Biomedical Platform
Remote Data Engineer for AI Biomedical Platform

Washington University in St. Louis • St. Louis (MO)

On-site
USD 75,000 - 129,000
Vacation & holidays
Health insurance
Transit pass
+1
Remote Data Engineer for AI Biomedical Platform
Remote Data Engineer for AI Biomedical Platform

Wustl • United States

Remote
USD 120,000 - 180,000
Remote work
Flexible schedule
C-BRAIN Data Engineer (Remote) - Neurology
C-BRAIN Data Engineer (Remote) - Neurology

Washington University • St. Louis (MO), Northern (KY)

Remote
USD 120,000 - 180,000
Data Engineer
Data Engineer

Laureate Institute for Brain Research • Tulsa (OK)

On-site
USD 110,000 - 150,000
Neural Data Infrastructure Engineer
Neural Data Infrastructure Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 180,000
Database Architect
Database Architect

Acro Service Corp • Saint Paul (MN)

On-site
USD 120,000 - 160,000
Neural Data Infrastructure Engineer
Neural Data Infrastructure Engineer

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 130,000 - 190,000
Research Software Engineer, Brain Data Science Platform (24-Month Fixed-Term)
Research Software Engineer, Brain Data Science Platform (24-Month Fixed-Term)

Stanford University • Palo Alto (CA)

On-site
USD 130,000 - 180,000