Duration : 12 - months contract (high possibility of extension)
Location : Cambridge MA
Job Requirement
We are seeking a highly skilled Data Engineer with expertise in Databricks, Seqera Platform (Nextflow Tower), cloud data engineering, and scientific data workflows to support our Discovery R&D and data science initiatives. The ideal candidate will design, develop, and maintain scalable data platforms and automated bioinformatics/data science pipelines that enable researchers and scientists to efficiently process, analyze, and access large-scale scientific and business datasets. This role requires strong experience in cloud-native architectures, data engineering best practices, workflow orchestration, and collaboration with cross-functional teams including scientists, bioinformaticians, data scientists, and IT infrastructure teams.
About the Role
The role involves designing, developing, and maintaining scalable data platforms and automated bioinformatics/data science pipelines.
Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies.
- Build and optimize ETL/ELT processes for structured, semi-structured, and unstructured data.
- Develop data ingestion frameworks for research, laboratory, clinical, and external scientific datasets.
- Implement data quality, validation, monitoring, and governance processes.
- Support enterprise data lakehouse architecture and data platform modernization initiatives.
- Databricks Administration & Development
- Develop and maintain Databricks notebooks, workflows, Delta Live Tables, and Jobs.
- Create optimized Spark-based transformations and data processing solutions.
- Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management.
- Manage Delta Lake environments and optimize performance, scalability, and cost.
- Integrate Databricks with cloud-native services and enterprise applications.
- Seqera Platform & Scientific Workflow Management
- Deploy, configure, and support Seqera Platform (formerly Nextflow Tower).
- Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads.
- Integrate Seqera workflows with AWS cloud infrastructure and compute environments.
- Support containerized workflows using Docker and Kubernetes technologies.
- Enable reproducible, scalable, and compliant scientific data processing workflows.
- Design and implement cloud-based data solutions in AWS.
- Manage cloud storage solutions including S3 and data lifecycle policies.
- Develop Infrastructure-as-Code solutions using Terraform or CloudFormation.
- Implement security controls and access management following enterprise IT standards.
- Partner with data scientists, researchers, bioinformaticians, and business stakeholders to understand data requirements.
- Provide technical guidance on data engineering best practices and workflow automation.
- Troubleshoot pipeline failures, performance issues, and workflow bottlenecks.
- Contribute to platform roadmaps and continuous improvement initiatives.
- Maintain technical documentation, SOPs, and knowledge articles.
Qualifications
Education
- Bachelor's degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field.
- Master's degree preferred.
Experience
- 5+ years of experience in data engineering, cloud engineering, or analytics platform development.
- 3+ years of hands-on experience with Databricks and Apache Spark.
- 2+ years of experience with Seqera Platform (Nextflow Tower) and Nextflow workflows.
- Experience supporting scientific research, life sciences, pharmaceutical, biotech, or healthcare environments preferred.
Required Skills
- Databricks & Data Engineering, Databricks Lakehouse Platform
- Apache Spark (PySpark, Spark SQL)
- Databricks Workflows
- Unity Catalog
- SQL and Python
- Seqera & Scientific Computing; Seqera Platform / Nextflow Tower
- Bioinformatics workflow automation
- Docker and container technologies
- High-performance computing environments
- S3, IAM, EC2, VPC, Lambda
- Terraform or CloudFormation
- Cloud monitoring and logging tools
- Data Technologies
- Data Lake and Lakehouse architectures
- ETL/ELT frameworks
- Data modeling
- Data cataloging and governance
- API integrations
- Data quality frameworks
- DevOps & Automation
- GitHub/GitLab
- Jenkins, GitHub Actions, or similar tools