We are currently hiring for an experienced Data Scientist / Data Engineer in Berlin for one of our clients which is a Berlin-based medical technology company.
The role focuses on building the data and AI infrastructure behind next-generation medical diagnostic solutions, working with high-frequency physiological and clinical data and supporting the transition of machine learning models from research into secure, scalable, production-ready clinical applications.
You will work closely with software, clinical, regulatory, and data science teams in a highly interdisciplinary environment.
This role is for someone who understands Data Science and the math behind it, Statistical Math and can work with a Senior Scientist in the team to understand the algorithms and build a production grade application from it.
Key Responsibilities
- Design, build, and maintain robust, scalable, and secure data ingestion and transformation pipelines.
- Implement automated data ingestion, validation, lineage tracking, and data quality monitoring.
- Develop reusable data libraries, transformation logic, feature engineering modules, and governance frameworks.
- Build data infrastructure for high-frequency physiological, sensor, biomedical, and clinical data.
- Ensure data platforms and pipelines comply with GDPR, NIS2, Cyber Resilience Act (EU), HIPAA (US), and other applicable regulatory requirements.
- Implement and maintain Role-Based Access Control (RBAC), encryption in transit and at rest, and secure authentication mechanisms.
- Develop data anonymization and pseudonymization workflows, consent handling, and audit mechanisms.
- Participate in security audits, compliance reviews, and certification activities.
- Operationalize AI and ML pipelines, including model deployment, monitoring, versioning, drift detection, and automated retraining.
- Integrate explainable AI tooling such as SHAP and LIME into ML pipelines to support auditability and regulatory compliance.
- Maintain CI/CD workflows for data and machine learning workloads.
- Develop Python-based analytics scripts, reusable data modules, and feature engineering logic.
- Support visualization of data and technical workflows using appropriate visualization and dashboarding tools.
- Ensure reproducibility, model traceability, and governance across the complete data and ML lifecycle.
- Collaborate with software, clinical, regulatory, and research teams to translate technical solutions into clinically relevant applications.
Required Skills / Primary Skills
- 5+ years of experience building secure and compliant cloud-based data ecosystems, preferably within regulated data domains.
- Hands-on experience working with time-series data, sensor data, biomedical data, physiological data, or wearable data.
- Strong experience designing and implementing scalable data ingestion and transformation pipelines.
- Advanced experience with data orchestration, data frameworks, and infrastructure automation.
- Strong understanding of machine learning approaches for classification, anomaly detection, and prediction using high-frequency data.
- Proficiency in Python or another programming language used for data engineering and machine learning.
- Strong understanding of statistical programming and data analysis techniques.
- Hands-on experience implementing data security and privacy mechanisms, including RBAC, encryption, anonymization/pseudonymization, consent handling, and auditability.
- Experience operationalizing machine learning models, including deployment, monitoring, versioning, drift handling, and retraining.
- Experience working with cloud-based data platforms and regulated data environments.
- Strong understanding of data quality, lineage, reproducibility, and model traceability.
- Ability to work effectively with software, clinical, regulatory, and data science teams.
- Clear, documentation-driven engineering approach.
- Bachelor's or Master's degree in Computer Science, Engineering, Data Science, AI, Cybersecurity, or a related field.
Additional Skills
- Experience working with multilevel longitudinal data, missing data strategies, and clinical outcome modeling.
- Good working knowledge of AWS, Azure, or GCP.
- Hands-on experience with Python and/or JavaScript.
- Experience with ML frameworks such as TensorFlow or PyTorch.
- Experience with visualization and monitoring tools such as Tableau, Looker, Grafana, or Streamlit.
- Experience with orchestration tools such as Airflow, Dagster, or Prefect.
- Experience with data frameworks such as Spark, dbt, or Kafka.
- Experience with infrastructure automation tools such as Terraform, Kubernetes, or Helm.
- Knowledge of medical data standards and formats such as FHIR, HL7, or DICOM.
- Experience managing encryption standards and secure authentication.
- Experience working in medical technology, healthcare, clinical research, or other highly regulated environments.
- Strong analytical and problem-solving skills.
- Excellent written and verbal communication skills.
Employment Details
- Location: Berlin, Germany
- Industry: Medical Technology / Healthcare
- Role: Data Scientist / Data Engineer
- Experience: 5+ years
If you're an experienced Data Scientist or Data Engineer with strong expertise in secure data engineering, high-frequency physiological data, cloud platforms, and production ML infrastructure, this is an opportunity to work at the intersection of medical devices, AI, data science, and clinical research and contribute to the development of next-generation med-tech solutions.