AI Data Architect / Senior Data Engineer

Infosys

Dadri

On-site

INR 3,500,000 - 6,000,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infosys is seeking an AI Data Architect to own the enterprise data pipeline powering an agentic AI platform. You will design ingestion, transformation, contextualisation, enrichment, validation, and semantic modelling layers that connect structured and unstructured data into an AI‑ready corpus.

This senior individual contributor role requires hands‑on experience with data engineering, ontology implementation, knowledge graphs, and cloud‑scale pipeline engineering, with responsibility for

Qualifications

  • 14+ years of experience in data engineering, data architecture, platform engineering, or enterprise‑scale data solution delivery.

Responsibilities

  • Design production‑grade enterprise connectors and ETL/ELT pipelines for structured and unstructured enterprise data sources.
  • Build ingestion and transformation pipelines using Python, SQL, PySpark, Apache Spark, Airflow, dbt, Dagster, Flink, or equivalent technologies.
  • Create frameworks for data labelling, contextualisation, harmonisation, enrichment, and classification workflows for AI agents.
  • Architect integration with knowledge graphs and vector databases for hybrid search and AI‑ready data access.
  • Build and maintain Ontology/knowledge graph pipelines using Neo4j, RDF/OWL, Apache Jena, Stardog, GraphDB, or equivalent technologies.
  • Implement graph validation frameworks such as SHACL or ShEx to enforce data integrity rules over enterprise knowledge graphs.
  • Implement data quality automation using Great Expectations, AWS Glue DataBrew, dbt tests, and other validation pipelines.
  • Define data profiling routines for real‑world enterprise data and data quality checks.
  • Collaborate with AI/ML architects to align pipelines with contracts, ontology models, and downstream AI consumption.
  • Mentor junior data engineers and help establish engineering practices for a robust data platform.

Skills

Databricks
Snowflake
BigQuery
Python
SQL
PySpark
Apache Spark
Data ingestion
Knowledge graphs
Neo4j
RDF/OWL
Graph databases
Vector databases

Tools

Airflow
dbt
Dagster
Flink
Pinecone
Weaviate
Milvus
Chroma
Neo4j
RDF/OWL

Job description

As an AI Data Architect, you will own the data pipeline that powers our agentic AI platform. You will design the ingestion, transformation, contextualization, enrichment, validation, and semantic modelling layers that connect a wide range of structured and unstructured enterprise data sources into an AI-ready data corpus.

This is a senior individual contributor role with real ownership over the data foundation of the venture. The role requires hands‑on experience in enterprise data engineering, schema discovery, data profiling, data quality automation, ontology implementation, knowledge graph integration, and cloud‑scale pipeline engineering.

Key Responsibilities
  • Design production‑grade enterprise connectors and ETL/ELT pipelines for both structured enterprise systems such as ERP, CRM, OSS/BSS, billing, finance, HR, and unstructured sources such as emails, documents, logs, transcripts, and media files.
  • Build ingestion and transformation pipelines using Python, SQL, PySpark, Apache Spark, Airflow, dbt, Dagster, Flink, or equivalent technologies.
  • Create frameworks for data labelling, contextualisation, harmonisation, enrichment, and classification workflows to configure AI agents.
  • Architect integration with knowledge graphs and vector databases for hybrid search, semantic retrieval, contextual reasoning, and AI‑ready data access.
  • Build and maintain Ontology/knowledge graph pipelines using Neo4j, RDF/OWL, Apache Jena, Stardog, GraphDB, or equivalent technologies.
  • Implement graph validation frameworks such as SHACL or ShEx to programmatically enforce data integrity rules over enterprise knowledge graphs.
  • Implement data quality automation using frameworks such as Great Expectations, AWS Glue DataBrew, dbt tests, custom validation pipelines, or equivalent tools.
  • Define data profiling routines for real‑world enterprise data, including missing keys, duplicate entities, inconsistent encoding, changing column meanings, incomplete master data, and conflicting source records.
  • Experience implementing semantic guardrails, jailbreak protection, data exfiltration prevention, and toxic output mitigation. RabbitMQ / Apache Kafka (Agent Message Queuing).
  • Implement privacy and compliance controls, including masking, anonymisation, access control, PII handling, GDPR compliance, and Indian Digital Personal Data Protection Act / DPDP Act alignment.
  • Partner with AI/ML architects to ensure pipeline outputs match agent input contracts, retrieval requirements, ontology models, and downstream AI consumption patterns.
  • Mentor junior data engineers, lead design reviews, and help establish engineering practices for a high‑quality, product‑grade data platform.
Must‑Have Qualifications
  • 14+ years of experience in data engineering, data architecture, platform engineering, or enterprise‑scale data solution delivery.
  • Strong production experience on modern data platforms such as Databricks, Snowflake, BigQuery, cloud data lakes, lakehouses, or equivalent enterprise data platforms.
  • Deep working knowledge of Python, SQL, PySpark, Apache Spark, and modern data pipeline development practices.
  • Hands‑on experience with both structured and unstructured data ingestion at enterprise scale.
  • Strong experience in building pipelines for enterprise sources such as ERP, CRM, OSS/BSS, billing systems, finance systems, ServiceNow, Salesforce, SAP, Oracle, and legacy databases.
  • Working knowledge of vector databases such as Pinecone, Weaviate, pgvector, Milvus, Chroma, or equivalent technologies.
  • Hands‑on knowledge of knowledge graphs, graph data modelling, graph querying, and enterprise graph implementation using Neo4j, Cypher, RDF, OWL, or equivalent technologies.
  • Experience with semantic data models, ontologies, industry data standards, or domain‑specific enterprise taxonomies.
Good to Have
  • Exposure to telecom, BFSI, manufacturing, or other complex enterprise domains.
  • Experience with OSS/BSS, ERP, CRM, billing, order management, product catalogue, service inventory, or network inventory systems.
  • Experience with RDF triple stores such as Apache Jena, Stardog, GraphDB, Amazon Neptune, or equivalent technologies.
  • Experience with data catalogues, metadata management tools, lineage platforms, or governance platforms.
  • Contributions to open‑source data tooling, graph tooling, ontology tooling, or data quality frameworks.
Why This Role Is Exciting

You will architect the data foundation from day one. Your designs will shape how the platform discovers, understands, contextualises, validates, and prepares enterprise data for AI consumption.

This role offers the opportunity to build the core data backbone for a venture‑backed Infosys platform, influence early product architecture, work closely with AI/ML architects, and create a scalable foundation that can support multiple industries and enterprise clients over time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineering AI Architect
Data Engineering AI Architect

Infosys • Bengaluru

On-site
INR 3,000,000 - 6,000,000
AI Data Engineer
AI Data Engineer

EXL • Gurugram District

On-site
INR 1,500,000 - 2,800,000
AI Data Platform Engineer / Architect
AI Data Platform Engineer / Architect

HCLTech • Dadri

On-site
INR 3,800,000 - 6,000,000
Data Architect - Senior Manager
Data Architect - Senior Manager

Ascendion • Maharashtra

On-site
INR 4,200,000 - 7,000,000
AI Data Architect
AI Data Architect

Tata Consultancy Services • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Lead Data Architect (AWS)
Lead Data Architect (AWS)

ANRGI TECH Pvt. Ltd. • Bengaluru Urban

On-site
INR 3,200,000 - 4,700,000
Data Architect
Data Architect

Laksh Human Resource • Dadri

On-site
INR 1,400,000 - 2,100,000
Agentic AI Data Engineer (Greater Noida)
Agentic AI Data Engineer (Greater Noida)

Kyndryl • Ghaziabad District

On-site
INR 1,800,000 - 3,000,000
Senior Data Engineer
Senior Data Engineer

Aon plc • India

Hybrid
INR 1,600,000 - 2,600,000
Required. AI & Data Analytics Engineer AI & Data Analytics Engineer
Required. AI & Data Analytics Engineer AI & Data Analytics Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

Hybrid
INR 2,500,000 - 4,500,000