AI Data Engineer

Agilent

España

Presencial

EUR 60.000 - 90.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Agilent seeks an AI Data Engineer to build domain data products and pipelines for model-ready consumption. You will ensure semantic annotation, data contracts, and quality scoring from day one, working with both on-premise and cloud data assets.

The role involves deploying agents for metadata generation, entity resolution, and content classification, and collaborating with data owners to ground the pod's retrieval foundations.

Formación

  • Strong data engineering with AI-focused data products.
  • Experience building retrieval-ready, semantically annotated data.
  • Familiarity with data contracts, lineage, and certification.

Responsabilidades

  • Develop domain data products and pipelines for AI consumption.
  • Ensure data quality and model-readiness with validation signals.
  • Collaborate with data owners and stewards to ground pod data.
  • Create grounding, retrieval, and storage foundations for the pod.
  • Deploy metadata generation and entity resolution agents.
  • Promote reusable assets and certify data products for enterprise use.

Conocimientos

Data engineering for AI
Retrieval-ready data
Semantic data contracts
Platform tech familiarity

Herramientas

Microsoft Fabric
Snowflake
Vector stores
Graph stores

Descripción del empleo

Agilent inspires and supports discoveries that advance the quality of life by providing life science, diagnostic and applied market laboratories worldwide with instruments, services, consumables, application and both measurement and asset management expertise.

Builds the data products and pipelines a pod runs on; turns raw domain data into model-ready assets. Pods do not wait for the Fabric to be complete; they build the Fabric through execution, and this role is where that happens for the data plane. Every domain data product built in a pod is constructed to certification standards from the start, because the second consumer of the asset is the point, not an afterthought.

The role goes beyond the conventional pipeline engineering. AI consumption changes what "model-ready" means: data must be retrieval-ready, semantically annotated, contract-governed, and quality-scored, and the AI Data Engineer often uses agents to do the building, generating metadata, resolving entities, and classifying unstructured domain content rather than hand-curating at a scale that cannot hold.

Responsible for
  • Domain data products and pipelines serving the pod's use case, built to the data plane's certification standards: semantic definition, data contract, entitlement metadata including agent identity, lineage, and certification tier from day one.
  • Data quality and model-readiness, including quality scoring against the defined dimensions and the quality signals that feed the evals spine; when quality falls below threshold, this role is the one who says so before the agent does something embarrassing with it.
  • Working relationships with the domain's data owners and stewards, so that steward-validated definitions and certified sources ground the pod's retrieval rather than whatever was easiest to reach.
  • Retrieval foundations for the pod: structured and unstructured grounding, vector and graph assets where the use case requires relationship reasoning, built on the platform estate rather than parallel infrastructure.
  • AI-built curation in practice: deploying metadata generation, entity resolution, and content classification agents against the domain's data, contributing those outputs to the registry.
  • Reusable assets back to the Fabric: every domain data product is a candidate for certification and enterprise reuse, documented and handed to the data plane for promotion, not a point integration that dies with the pod.
  • Works with the business teams and IT to develop domain data models.
  • Applying broad understanding of on premise and cloud deployment topologies, creates data collection frameworks to capture, manage, store and utilize large sets of structured, unstructured and/or disconnected data from a wide variety of internal and external sources.
  • Oversees the establishment of data set processes and builds data structures based on business and technical requirements to funnel data from multiple sources and store using any combination of storage forms (e.g. cloud, local databases) as required.
  • Consults in design standards and assurance processes for software, systems and applications development to ensure compatibility and operability of data connections, flows and storage requirements.
  • Creates data tools for analytics and data scientist team members that assist them in building and optimizing our analytics products.
  • Prepares and manipulates data for predictive and prescriptive modeling. Uses data to discover tasks that can be automated. Identify ways to improve data reliability, efficiency and quality.
  • Reviews internal and external business and product requirements for data operations and activity and suggests changes and upgrades to systems and storage to accommodate ongoing needs.
What success looks like in year one
  • The pod's use case running entirely on contract-governed, quality-scored data products, with no undocumented side channels into source systems.
  • Multiple domain data products from the pod certified into the registry and consumed or queued for consumption by multiple use cases.
  • Quality signals from the pod's domain flowing into the evals spine, with at least one instance where a quality threshold correctly gated an agent behavior.
  • Measurably faster data-to-build time driven by reuse and process optimization
Qualifications
  • Strong data engineering with experience building for AI consumption: retrieval-ready, semantically annotated, contract-governed data products, not only warehouse tables and dashboards.
  • Hands-on familiarity with the platform estate (Microsoft Fabric, Snowflake, vector and graph stores) and with operating under data contracts, lineage, and certification requirements.
  • Experience with RAG data foundations: chunking, embedding, hybrid retrieval, and the failure modes that show up in agent behavior rather than in pipeline monitoring.
  • The disposition to work inside a business domain: you interview stewards, read the pipeline code that produces the data to understand what it act
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Enterprise AI Product Lead
Enterprise AI Product Lead

Agilent • España

Presencial
EUR 90.000 - 130.000
AI Data Engineer
AI Data Engineer

Agilent Technologies Spain S.L. • Bellprat

Presencial
EUR 149.000 - 233.000
AI Data Engineer: Build Model-Ready Data Pipelines
AI Data Engineer: Build Model-Ready Data Pipelines

Agilent • España

Presencial
EUR 60.000 - 90.000
Enterprise AI Product Lead
Enterprise AI Product Lead

Agilent Technologies • Barcelona

Presencial
EUR 90.000 - 120.000
AI Harness Engineer
AI Harness Engineer

Agilent Technologies • Barcelona

Presencial
EUR 90.000 - 120.000
Equal opportunity employer
AI Analyst (UA/RU Language speaking)
AI Analyst (UA/RU Language speaking)

Neurons Lab • España

Presencial
EUR 60.000 - 100.000
AI Software Engineer | Spain
AI Software Engineer | Spain

Accenture España • Madrid

Presencial
EUR 90.000 - 130.000
AI Operations Engineer
AI Operations Engineer

United States Digital Space LLC • Amer

Presencial
EUR 70.000 - 95.000
AI Engineer
AI Engineer

Werfen • Barcelona

Presencial
EUR 50.000 - 70.000
Forward Deployed Manager
Forward Deployed Manager

Accenture España • Madrid

Presencial
EUR 90.000 - 150.000