Senior ML Data Processing Developer - AI Safety Research Non-Profit - Berlin (Hybrid)
We are seeking a Senior ML Data Processing Developer to contribute to the development, curation and scaling of large-scale machine learning data pipelines. In this role, you will go beyond traditional data engineering by actively engineering data quality. You will design algorithmic filtering systems, develop model-based scoring mechanisms, implement rigorous data-quality controls and build novel data transformations for emerging machine learning requirements.
Key Responsibilities
- Partner with Research and Engineering teams to define, build, automate, scale and maintain data pipelines that transform web-scale data into high-quality training datasets.
- Build and maintain large-scale data processing pipelines covering deduplication, model-based quality scoring, heuristic filtering, toxicity removal, PII scrubbing, metadata extraction and custom data transformations.
- Ensure datasets have robust versioning, lineage and provenance tracking while optimising pipelines for throughput and cost.
- Ensure ingested and processed data meets relevant compliance requirements, internal data governance policies and legal obligations.
- Develop and refine data-quality tooling, including:
- Heuristic filtering systems
- LLM-as-a-judge evaluators
- Machine learning classifiers
Skills & Qualifications
- Degree in computer science, software engineering, or a related field.
- Proven track record of handling massive unstructured text datasets (trillion-token scale), with 5+ years of experience in data processing, machine learning engineering or Natural Language Processing (NLP).
- Hands-on experience with distributed processing frameworks (e.g., Spark, Ray, Flink), designing and optimizing high-throughput pipelines.
- Experience with data privacy implementation (PII scrubbing), content-safety filtering (toxicity, bias), and evaluation-contamination prevention.
- Demonstrated ability to work across Research, Engineering and/or Legal/Governance teams, translating varied requirements into concrete pipeline work.
- Strong Python proficiency, including experience writing production-grade data-processing code.
- Experience with pipeline orchestration frameworks (e.g., Airflow, Prefect, Dagster).
Nice to have
- Experience training, fine-tuning, or deploying ML models for data-quality tasks (classifiers, LLM-based evaluators) and familiarity with LLM inference optimization (e.g. vLLM, SGLang).
- Familiarity with containerized deployment (Docker, Kubernetes) and infrastructure-as-code practices.
- Familiarity with ML experiment tracking tools (e.g. Weights and Biases).
- Experience with data licensing workflows or web-scale data acquisition.
- Contributions to open-source data processing or NLP tooling.