Turn this role into an interview — a resume and cover letter built around what this employer wants.
Stanford University’s Tolias Lab in Ophthalmology seeks a Data Engineer to design, build, and operate end-to-end data pipelines from acquisition through ETL/ELT, enabling research workflows and scalable data platforms. You will partner with scientists and engineers to implement interfaces, observability, and dependable tools, ensuring robust data movement and governance across the lakehouse stack.
The role emphasizes productionizing research code, building durable data models, and collaborating
The Tolias Lab (https://toliaslab.org/) in the Department of Ophthalmology (http://med.stanford.edu/ophthalmology.html) in the School of Medicine is seeking a Data Engineer (Software Developer 2) to build and maintain the data platform supporting Enigma’s experimental workflows. You will partner with scientists and engineers to turn evolving experimental requirements into clear interfaces, observable pipelines, and dependable tools. You will develop reliable systems that move raw data from data acquisition systems through ETL/ELT transformation into downstream data and metadata stores. This role will involve data lakehouse and schema design and administration, data governance and observability, and performance optimization.
The ideal candidate enjoys owning systems end to end while working in a highly collaborative research environment.
Enigma is building the experimental and computational infrastructure needed to understand intelligence through large-scale neuroscience. Our teams work across experimental development, neurophysiology, data engineering, and machine learning to turn complex biological data into reliable, accessible scientific resources.
Within Enigma, the Experimental Development (xDev) team designs and supports neurophysiology and behavioral experiments. xDev develops the systems that connect experimental hardware and acquisition software with data processing and downstream analysis—working closely with our data infrastructure engineering and our scientific data analysis and modeling teams.
Design, build, and operate ETL/ELT and high-throughput ingestion pipelines connecting acquisition software, processing code, metadata registries, object storage, and databases
Own the data layer end-to-end: schema design, indexing, query optimization, migrations, backups, and durable data models that hold up over years of scientific work
Improve reliability, scalability, observability, and performance across the data stack, and lead troubleshooting/incident response when things break
Build tools with and to help researchers and our in-house AI agents discover datasets, inspect processing state, and safely diagnose or rerun failed jobs
Productionize research code — testing, packaging, deployment, monitoring, documentation
Partner with data acquisition, infrastructure engineering, scientific analysis and AI modeling teams.
Strong proficiency in Python and SQL
Experience with relational and NoSQL databases, as well as object storage (e.g., S3, MinIO, Ceph)
Experience designing, building, and operating production ETL/ELT pipelines
Experience designing durable data models and schemas for complex, evolving datasets
Experience building high-throughput ingestion systems and optimizing data movement across storage, compute, and network boundaries
Experience with Docker and container orchestration platforms such as Kubernetes
Strong communication skills and experience working in cross-functional teams
Experience designing APIs or service interfaces used by multiple teams
Experience with scientific computing and large-scale neuroscience or other scientific data
Workflow orchestration experience with Airflow, Dagster, Prefect, or similar tools
Experience in building pipelines with and for autonomous, agentic AI systems
Experience with performance-critical compiled or systems languages (C, C++, or Rust)
Familiarity with data provenance and lineage tracking, metadata catalogs, or dataset versioning
Are excited by a fast-paced, production-focused research environment that often requires switching between many hats
Can move quickly without sacrificing rigor
Care about making complex workflows reliable and easy to use
Enjoy translating ambiguous research needs into pragmatic technical systems
Take ownership across design, implementation, deployment, and operations
Enjoy collaborative design and working closely across team boundaries
Are energized by working as part of a cohesive team pursuing ambitious, long-term neuroscientific and AI goals
A highly collaborative environment across neuroscience, engineering, and AI
Opportunity to contribute to a next-generation neurotechnology platform
Competitive salary and benefits
Strong mentoring and career development support