Data Engineer for AI (f/m/d)

Allianz

Frankfurt

Vor Ort

EUR 70.000 - 110.000

Vollzeit

Vor 10 Tagen
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Allianz Global Investors is seeking a Data Engineer for AI to design, build and operate data products on the Databricks-based Data and AI Platform in Frankfurt. You will enable platform customers to source, ingest, transform and serve high-quality data for AI/ML use cases with self-service access and robust governance.

You will collaborate with ML/AI engineers and platform architects to evolve ingestion, transformation, and governance features, ensuring compliance and scalable, reusable data

Qualifikationen

  • 3+ years building and operating data pipelines on Databricks (batch and streaming) with governance.
  • Expert-level Databricks, Spark, Python and SQL, plus CI/CD and Terraform on Azure.
  • Diploma in computer science or a similar field.
  • Strong data architecture knowledge and governance awareness.
  • Excellent English communication skills.

Aufgaben

  • Develop and maintain standardized pipeline patterns for AI/ML-ready datasets.
  • Build robust batch and streaming pipelines on Databricks to ingest data from internal and external sources.
  • Operate curated data products and feature store with proper versioning and lightweight documentation.
  • Collaborate with ML/AI engineers to build GenAI/RAG pipelines and data ingestion workflows.
  • Provide L3 support for pipeline and data product incidents with root-cause analysis.
  • Ensure security, privacy and regulatory compliance via Unity Catalog and governance controls.
  • Optimize Spark workloads and storage layouts for scalable, cost-efficient pipelines.
  • Partner with data scientists and DevOps to enable self-service through documentation and reviews.

Jobbeschreibung

In the role of Data Engineer for AI (f/m/d), you will design, build and operate data products and reusable data preparation components on the Databricks-based Data and AI Platform at Allianz Global Investors.

Your focus is to enable platform customers to reliably source, ingest, transform, validate and serve high-quality, compliant data for AI and ML use cases (including analytics and GenAI), so teams can consume data in a self-service manner. You will provide technical guidance and L3 support to delivery teams, define and promote best practices for data pipelines and data quality, and ensure adherence to security, privacy and regulatory requirements following a compliance-by-design approach.

In addition, you will continuously evolve the platform’s data engineering capabilities by integrating new features (e.g., ingestion patterns, transformation frameworks, governance controls and monitoring) and by delivering standardized, reusable pipelines and templates that scale across use cases.

This position will be based in Frankfurt.

What you will do
  • Develop, maintain and enhance standardized pipeline patterns, templates and utilities (ingestion, transformation, validation, enrichment) to deliver AI/ML‑ready datasets
  • Build robust pipelines for structured and unstructured data with reproducible, deterministic outputs
  • Build and operate reliable batch and streaming pipelines on Databricks to ingest data from internal and external sources, curate datasets, and publish trusted data products for AI, ML and analytical consumption
  • Curated data products (Bronze/Silver/Gold): Establish controlled dataset evolution, enforce data quality and freshness, and curate gold tables explicitly designed for analytical, ML and AI use cases
  • Design, build and operate feature engineering pipelines and curate a ready‑to‑use feature store with proper versioning and lightweight documentation (ownership, purpose, inputs/outputs, SLAs)
  • Collaborate with ML and AI Engineers to build and operate pipelines for GenAI/RAG systems, including new data ingestion, Bronze/Silver updates, chunking, embedding refresh and index updates to maintain knowledge base relevance
  • Enable systematic experimentation (offline evaluations, A/B tests) and support model retraining using user feedback (binary and non‑binary) to continuously improve models
  • Collaborate with the Databricks Platform Architect to introduce and operationalize platform capabilities for ingestion, transformation and serving (Delta/Unity Catalog patterns, orchestration, monitoring), making them available securely and consistently
  • Facilitate provisioning and rollout of Databricks Lakeflow connectors to integrate new data sources and standardize ingestion across domains
  • Implement and enforce governance using Unity Catalog, including access control, lineage, retention, encryption and auditing, ensuring compliance with internal standards and regulations
  • Establish automated data quality checks and validations (e.g., DQX), define SLAs/SLOs for freshness and reliability, and implement monitoring and observability
  • Optimize Spark workloads, storage layouts and compute usage to ensure scalable, stable and cost‑efficient data pipelines
  • Partner with Data Scientists, ML Engineers, AI CoE, DevOps and SecOps to define data requirements, promote reusable patterns, and enable self‑service through documentation, coaching and reviews
  • Provide L3 support for pipeline and data product incidents, perform root‑cause analysis, and drive continuous improvements to platform stability and customer satisfaction
  • Integrate Databricks with (Azure) cloud services, enterprise systems, external providers and SaaS platforms, and expose curated data products securely to downstream consumers (feature stores, model training/inference, BI)
What you bring
Required:
  • Minimum 3 years of practical experience building and operating data pipelines and data products at scale on Databricks (batch and/or streaming), ideally in environments with strong governance requirements
  • Strong understanding of data engineering concepts for AI/ML (data modelling, data quality, data contracts, feature engineering collaboration, reproducibility) and how data characteristics impact model outcomes
  • Expert‑level technology skills in Databricks, Spark, Python and SQL, plus solid engineering practices such as CI/CD and infrastructure‑as‑code (e.g., Terraform) on Azure
  • Diploma in computer science or a similar field
  • Very good understanding of data architecture and platform patterns (Lakehouse concepts, Delta, medallion approaches, data product thinking) and how to operationalize them in a governed enterprise context
  • Excellent analytical, planning, and organizational skills
  • Good oral and written communication skills in English
  • Ability to efficiently and effectively document and share knowledge and enable others
Preferred:
  • Working experience in the financial industry, preferably in asset management
  • International work experience
  • Certification “Databricks Certified Data Engineer Professional” (or equivalent)
  • Certification “Databricks Certified Data Engineer Associate” (or equivalent)
  • Optional: Certification “Databricks Certified Generative AI Engineer Associate” (helpful for GenAI‑related data preparation patterns)
What we offer
  • We empower our employees by ensuring flexible work arrangements that maintain a balance between performance, productivity, career development and personal priorities (e.g., hybrid model/ flexible working hours)
  • Securing your future: Access to company pension/savings plans
  • Family support (relocation/ childcare facilities)
  • Company share purchasing plan
  • Mental health and wellbeing programs
  • Mobility solutions (Jobrad bike leasing, subvention Jobticket)
  • Career opportunities within the entire Allianz Group
  • Self‑guided learning & development
  • Volunteering time
  • … and so much more!
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Engineer (f/m/d)
AI Engineer (f/m/d)

Allianz • Frankfurt

Vor Ort
EUR 90.000 - 130.000
Hybrid model / flexible working hours
Pension plans
Relocation support
+6
Senior Data Engineer – Databricks (m/w/d)
Senior Data Engineer – Databricks (m/w/d)

BRINK Group • München

Hybrid
EUR 90.000 - 130.000
30 days annual leave
Performance-based bonus
Workation up to 30 days/year
+2
Senior Backend Engineer (m/f/d)
Senior Backend Engineer (m/f/d)

Allianz Partners • München

Vor Ort
EUR 70.000 - 90.000
Training and development programs
Work-life balance initiatives
Senior Software Engineer (f/m/d) - Databricks Data Engineering & Analytics (part-/full-time)
Senior Software Engineer (f/m/d) - Databricks Data Engineering & Analytics (part-/full-time)

European Commodity Clearing AG • Leipzig

Hybrid
EUR 90.000 - 120.000
Childcare assistance
Meal allowance
Job ticket
+1
Platform Engineer (Cloud Infrastructure, AI Platform) (m/f/d)
Platform Engineer (Cloud Infrastructure, AI Platform) (m/f/d)

Allianz Partners • München

Vor Ort
EUR 65.000 - 85.000
Work Well programs
Career development courses
International mobility opportunities
Senior Quant Engineer Investment Platform (f/m/d)
Senior Quant Engineer Investment Platform (f/m/d)

Allianz Global Investors GmbH • Frankfurt

Hybrid
EUR 90.000 - 130.000
Hybrid work model
Company pension/savings plans
Relocation/ childcare facilities
+3
Cloud Data & AI Engineer (m/w/d) - Databricks
Cloud Data & AI Engineer (m/w/d) - Databricks

ruhrdot. • Essen

Hybrid
EUR 60.000 - 80.000
Zertifizierungen werden unterstützt
Flexibilität und Home-Office-Optionen
Best-in-Class Setup und Premium-Equipment
Data Engineer (Data Bricks)
Data Engineer (Data Bricks)

DataEdge SRL • Frankfurt

Vor Ort
EUR 85.000 - 110.000
AI Engineer (m/f/d)
AI Engineer (m/f/d)

Allianz Technology • Unterföhring

Hybrid
EUR 70.000 - 110.000
Company bonus scheme
Pension
Employee shares program
+1
Strategic Core Account Executive - Banking, m/f/d
Strategic Core Account Executive - Banking, m/f/d

Databricks Inc. • München

Vor Ort
EUR 90.000 - 150.000