European Tech Recruit are working closely with an exciting AI company committed to advancing research and creating technical solutions that enable safe-by-design AI systems.
They are looking for a talented Senior Data Platform Engineer to join them on a hybrid basis in Berlin.
In this high-impact role, you will bridge the gap between cutting-edge AI research and high-performance engineering, treating the data platform as an internal product with research teams as your primary customers.
You will be responsible for building automated, petabyte-scale data processing pipelines and ensuring they are efficient, reliable, reproducible and fully traceable.
Responsibilities as Senior Data Platform Engineer:
- Scale and automate data processing infrastructure to handle petabytes of data while ensuring reliable day-to-day operation.
- Design execution layers for pipeline stages with different computational requirements, selecting appropriate processing engines and maintaining efficient end-to-end throughput.
- Optimise the use of compute resources across large-scale workloads, including GPU resources for compute-intensive data processing tasks.
- Build reliable pipelines with graceful failure recovery, restartability and comprehensive observability.
- Identify and diagnose performance bottlenecks, attributing compute usage and costs to individual pipeline stages.
- Establish robust systems for dataset versioning, lineage, reproducibility and traceability across all stages of the data processing lifecycle.
- Ensure intermediate outputs and final datasets can be reliably reproduced for specific and evolving research requirements.
- Develop and maintain appropriate documentation and data specifications in line with internal data governance requirements.
- Partner closely with Research and Engineering teams to ensure datasets integrate seamlessly with model training pipelines.
- Evaluate and introduce new technologies and approaches where they can improve scalability, reliability, performance or cost efficiency.
Requirements:
- A bachelor's degree in a relevant field (e.g., computer science, computer engineering, software engineering) is required.
- 5+ years of experience designing, implementing, and managing large-scale distributed data processing systems, or working within large-scale distributed ML data frameworks, with recent experience using e.g. Ray, Apache Spark, workflow orchestrators, Apache Arrow, and/or Parquet.
- Demonstrated ownership of a data processing system under real throughput, reliability, and cost pressure. The specific frameworks matter less to us than evidence that you have had to reason about where a large pipeline breaks and why.
- Experience profiling and optimizing throughput and cost across heterogeneous workloads, including GPU-accelerated stages.
- Experience with dataset versioning, lineage, and reproducibility tooling.
- Ability to collaborate effectively with cross-functional teams, document best practices, and stay updated with the latest advancements in large-scale data processing and software development.
- Experience with workload managers (e.g., Ray, Kubernetes, Slurm).
- Familiarity with containerization tools (e.g., Docker, Enroot).
- Familiarity with data infrastructures and platforms (e.g., vector databases).
By applying to this role you understand that we may collect your personal data and store and process it on our systems. For more information please see our Privacy Notice (https://eu-recruit.com/about-us/privacy-notice/)