A complete application in a minute — tailored resume and cover letter, ready to send.
Accenture in India seeks a hands-on engineer to lead the development of a Semantic Knowledge Graph platform across a large data governance estate. You will build and maintain metadata extraction pipelines from Collibra, transform data for a Neptune-based graph, and create AI agent components for governed, intelligent access to graph knowledge.
The role requires 6 years of data engineering/AI experience, proficiency with graph databases, Collibra APIs, and Python, along with exposure to LLM-based
Packaged/SaaS App Engineering Lead
Own the design and evolution of packaged or SaaS solutions, overseeing configuration, integrations, and releases. Set platform standards to ensure stability, performance, and alignment with business and vendor constraints.
Generative AI
NA
15 years full time education
This position is an hands-on engineer role within a strategic enterprise initiative building a Semantic Knowledge Graph platform across a large data governance estate The incumbent will develop and maintain metadata extraction pipelines from the enterprise data catalog platform, transform and load metadata into a cloud-native graph database, and construct AI agent components that enable intelligent, governed consumption of graph knowledge.
Develop and maintain metadata extraction pipelines from Collibra via REST API, bulk export, delta detection, and event-triggered ingestion, handling both structured catalog objects and unstructured textual content.
Implement graph transformation logic — mapping extracted Collibra metadata to the program graph data model, resolving entity references, handling data quality anomalies, and producing Neptune-ready serialisations.
Execute and iterate on Neptune load procedures including bulk load, streaming upsert, and incremental refresh validate graph integrity post-load and maintain comprehensive load logs.
Build and test AI agent components under the direction of the Senior Manager: tool definitions, prompt templates, graph query generation, retrieval chains, and response synthesis modules.
Implement hybrid vector-graph retrieval patterns using embedding models and vector stores in conjunction with Neptune graph queries.
Write unit and integration tests for pipeline and agent components contribute to CI/CD pipeline configuration and maintenance.
Participate in Business Analyst collaboration sessions — capturing clarifications as structured requirements and translating use case feedback into concrete engineering changes.
Maintain technical documentation covering pipeline data flow diagrams, graph schema definitions, agent tool inventories, and known issue logs.
Monitor pipeline and agent behaviour in deployed environments investigate anomalies and contribute structured findings to root-cause analysis.
Support stakeholder demonstrations and validation sessions by preparing test scenarios, sample queries, and representative agent responses.
6 years of hands-on experience in data engineering or AI and machine learning engineering.
Working experience with at least one graph database — Amazon Neptune, Neo4j, or equivalent — including data modelling and query writing production exposure is preferred.
Practical experience with Collibra or a comparable enterprise data catalog platform, including programmatic API-based interaction.
Demonstrable experience building or contributing to LLM-based applications — RAG pipelines, agent frameworks, or tool-calling systems using LangChain, CrewAI, LangGraph, or equivalent frameworks.
Solid Python engineering skills with the ability to write clean, testable, and peer-reviewable code.
Familiarity with Gremlin or SPARQL at a functional working level.
AWS fundamentals: S3, Lambda, Glue, and introductory exposure to Neptune or OpenSearch.
Exposure to Healthcare or Life Sciences data environments — clinical data dictionaries, regulatory data structures, or pharma commercial data — at any level of depth genuine domain interest is required.
Exposure to ontology concepts: OWL, SKOS, RDF, or semantic data modelling principles.
Experience with embedding models and vector stores in a retrieval-augmented generation context.
Familiarity with Collibra business glossary, data lineage, or data quality objects at a module level.
Understanding of data governance principles: stewardship, data quality dimensions, and metadata lifecycle management.
Prior experience working within an agile delivery team operating Business Analyst-led requirements cycles.