Senior Data Platform / Data Engineer

Straumann Group

Warszawa

On-site

PLN 180,000 - 300,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Straumann Group is seeking a Senior Data Platform / Data Engineer to join our ML Platform team and help build the data infrastructure powering our AI products in dentistry. You will work with ML researchers, MLOps engineers, and product teams to ensure a reliable, scalable platform that supports the full AI lifecycle.

A key focus is the Data Lakehouse and dataset management workflows, including dataset versioning with DVC, and improving how data is prepared, extracted, and consumed across

Qualifications

  • Strong Python engineering skills.
  • Experience building data pipelines or data platforms.
  • Experience working with AWS.
  • Experience working with large datasets used in ML workflows.
  • Strong software engineering practices (testing, CI/CD, documentation).
  • Experience collaborating with ML teams or working in AI environments.

Responsibilities

  • Design and evolve the Data Lakehouse architecture used across our ML teams.
  • Improve the reliability and structure of data ingestion, extraction, and transformation pipelines.
  • Ensure datasets used for training and evaluation are consistent, reproducible, and well documented.
  • Collaborate with ML researchers to translate requirements into robust platform capabilities.

Skills

Python
Data pipelines
AWS
Large ML datasets
CI/CD
ML collaboration

Tools

DVC
Kubernetes
PostgreSQL
Metabase

Job description

Straumann Group At Straumann Group we’re on an exciting journey of growth, innovation, and impact - driven by our mission to improve oral health and transform millions of lives worldwide. United by purpose, we bring our best selves to work every day, embracing a high-performance, player-learner culture that inspires collaboration, curiosity, and ambition. Here, you’ll have the opportunity to take charge of your own career, harnessing your skills, passion, and enthusiasm for learning to continually grow and progress. Together, we’re not just shaping brighter smiles, we’re unlocking the potential of people everywhere, including our own.

Straumann Group At Straumann Group we’re on an exciting journey of growth, innovation, and impact - driven by our mission to improve oral health and transform millions of lives worldwide. United by purpose, we bring our best selves to work every day, embracing a high-performance, player-learner culture that inspires collaboration, curiosity, and ambition. Here, you’ll have the opportunity to take charge of your own career, harnessing your skills, passion, and enthusiasm for learning to continually grow and progress. Together, we’re not just shaping brighter smiles, we’re unlocking the potential of people everywhere, including our own.

About The Role

We are looking for a Senior Data Platform / Data Engineer to join our ML Platform team and help build and scale the data infrastructure that powers our AI products in dentistry. Our platform supports the full AI development lifecycle, from raw data ingestion and annotation workflows to dataset versioning and model training pipelines. You will work closely with Machine Learning Researchers (MLRs), MLOps engineers, and product teams to ensure our data infrastructure is reliable, scalable, and easy to use.

A key focus of the role is improving our Data Lakehouse (DLH) and dataset management workflows, including dataset versioning (DVC) and improving how data is prepared, extracted, and consumed across research and production systems.

What You Will Work On

You will play a key role in shaping the next generation of our data platform.

Typical Responsibilities Include

Data platform ownership

  • Design and evolve the Data Lakehouse (DLH) architecture used across our ML teams.
  • Improve the reliability and structure of data ingestion, extraction, and transformation pipelines.
  • Ensure datasets used for training and evaluation are consistent, reproducible, and well documented.

Dataset lifecycle management

  • Improve workflows for dataset versioning and reproducibility using tools such as DVC.
  • Design solutions for managing multiple versions of datasets and annotations across experiments and models.
  • Improve the ability for researchers to retrieve the correct dataset versions reliably.

Data pipelines and infrastructure

  • Build and maintain scalable data pipelines in Python.
  • Improve metadata management, dataset validation, and data quality monitoring.
  • Optimize data workflows across AWS-based infrastructure.

Collaboration with ML teams

  • Work closely with ML researchers and ML engineers to understand their data needs.
  • Support research workflows with reliable and efficient data access patterns.
  • Help translate research requirements into robust platform capabilities.>

Data governance and quality

  • Implement practices for data quality, reproducibility, and traceability across the ML lifecycle.
  • Ensure our data infrastructure meets the requirements of regulated AI development.
Must Have
What we’re looking for:
  • Strong Python engineering skills
  • Experience building data pipelines or data platforms
  • Experience working with AWS
  • Experience working with large datasets used in ML workflows
  • Strong software engineering practices (testing, CI/CD, documentation)
  • Experience collaborating with ML teams or working in AI environments
Nice To Have
  • Experience with dataset versioning tools such as DVC
  • Experience with Kubernetes
  • Experience with data lakehouse architectures
  • Experience working with annotation pipelines or ML training datasets
  • Experience with PostgreSQL, Metabase, or similar data tooling
  • Experience working in regulated environments (medical / healthcare AI)
Our stack
  • AWS
  • Python
  • Kubernetes
  • PostgreSQL
  • Metabase
  • DVC for dataset versioning
  • Internal Data Lakehouse infrastructure

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or disability.

Employment Type: Full Time

Alternative Locations: Spain : Madrid || Poland : Gdansk || Poland : Warsaw || Poland : Wroclaw

Travel Percentage: 0 - 10%

Requisition ID: 20071

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Platform / Data Engineer
Senior Data Platform / Data Engineer

Straumann Group • Wrocław

On-site
PLN 180,000 - 260,000
Senior Data Platform / Data Engineer
Senior Data Platform / Data Engineer

Straumann Group • Województwo pomorskie

On-site
PLN 180,000 - 280,000
Senior Data Platform Engineer - ML & Data Lakehouse
Senior Data Platform Engineer - ML & Data Lakehouse

Straumann Group • Warszawa

On-site
PLN 180,000 - 300,000
Senior Data Platform Engineer for AI/ML Pipelines
Senior Data Platform Engineer for AI/ML Pipelines

Straumann Group • Wrocław

On-site
PLN 180,000 - 260,000
Senior Data Platform Engineer - ML Pipelines & Lakehouse
Senior Data Platform Engineer - ML Pipelines & Lakehouse

Straumann Group • Województwo pomorskie

On-site
PLN 180,000 - 280,000
Data Intelligence Platform Engineer
Data Intelligence Platform Engineer

Philip Morris International • Kraków

On-site
PLN 179,000 - 214,000
Data Engineer
Data Engineer

Philip Morris International • Kraków

Hybrid
PLN 144,000 - 172,000
Life and Health insurance
Employee Pension Plan
Hybrid working model
+1
Senior Data Scientist
Senior Data Scientist

HEDONE • Województwo pomorskie

Remote
Salary 25,000 - 33,000 PLN+ VAT (B2B) or 25,000 - 33,000 PLN gross (mandate)
Paid Holidays
Travel to California
Senior Data Scientist
Senior Data Scientist

SKMgroup • Województwo małopolskie

On-site
PLN 60,000 - 90,000
Large freedom and real influence
Team approach to challenges
Flexible working culture
+1
Data and AI Engineer - RDT Pharma R&D
Data and AI Engineer - RDT Pharma R&D

Roche • Warszawa

On-site
PLN 229,000 - 425,000