RAG Architect

Zoho

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Zoho in Bengaluru seeks an experienced Context Window Optimization / RAG Architect to own enterprise generative AI retrieval performance. You will design high-throughput knowledge retrieval systems, optimize semantic context parsing, and build re-ranking and caching pipelines to deliver accurate, data-grounded responses with minimal latency.

The role focuses on end-to-end RAG pipelines, context window strategies, and hybrid search algorithms, requiring 6–10 years in data engineering and 3+ years

Qualifications

  • - 6 to 10 years of enterprise data engineering, database design, or search-engine engineering experience.
  • - 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
  • - Strong mastery of Python, vector databases, text embedding models, and open-source orchestration tools (LlamaIndex, LangChain).
  • - Mandatory cloud data/database engineer or specialty analytics certification from a major cloud vendor (AWS/GCP/Azure).

Responsibilities

  • Architect end-to-end Retrieval-Augmented Generation (RAG) pipelines with document parsing and multi-vector lookups.
  • Optimize context window utilization patterns with smart chunking and sliding window strategies.
  • Build high-performance re-ranking layers using cross-encoders to score retrieved documents.
  • Implement automated semantic caching architectures to reduce API token costs and latency.
  • Establish automated data chunking pipelines for semi-structured and unstructured inputs into vector targets.
  • Govern vector similarity spaces by combining dense embeddings with sparse BM25 indexes.
  • Audit context-level hallucination rates and accuracy logs to improve retrieval performance.

Skills

Python
Vector databases
Text embedding models
SQL
LangChain
LlamaIndex
BM25 / vector search
Cloud data engineering

Tools

Pinecone
Milvus
Weaviate
Neo4j

Job description

  • Total Experience Required: 6 to 10 years
  • Relevant Experience Required: 3+ years of dedicated experience designing production-grade Retrieval-Augmented Generation (RAG) architectures and optimizing LLM token throughput
Job Summary

We are seeking an experienced Context Window Optimization / RAG Architect to take full ownership of our enterprise generative AI retrieval performance, accuracy, and operational cost metrics. The ideal candidate will design high-throughput knowledge retrieval systems, optimize semantic context parsing, build custom re-ranking pipelines, and engineer caching grids to deliver data-grounded AI responses with minimal latency and maximum token efficiency.

Key Responsibilities

  • Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, building structures for document parsing, semantic metadata enrichment, and multi-vector lookups.
  • Optimize context window utilization patterns, designing smart parent-child chunking models, sentence-window retrievals, and sliding window strategies to eliminate irrelevant text tokens.
  • Build high-performance re-ranking layers, deploying machine learning cross-encoders (e.g., Cohere Rerank, BGE-Reranker) to score retrieved documents before feeding them into the LLM context pool.
  • Implement automated semantic caching architectures, utilizing caching layers (e.g., GPTCache) to capture recurring semantic queries, reducing API token expenditures and response latencies.
  • Establish automated data chunking pipelines, configuring ingestion routines to cleanly parse semi-structured and unstructured formats (PDFs, corporate wikis, SQL outputs) into clean vector targets.
  • Govern vector similarity spaces, fine-tuning hybrid search algorithms that cleanly combine dense semantic embeddings with sparse keyword token indexes (BM25).
  • Audit context-level hallucination rates and accuracy logs, tracking precision metrics, retrieval recall bounds, and processing speeds to systematically eliminate incorrect model generations.
Requirements
  • 6 to 10 years of enterprise data engineering, database design, or search engine engineering experience, with 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
  • Strong technical mastery of Python, vector databases (Pinecone, Milvus , Weaviate ), text embedding models, open-source orchestration tools ( LlamaIndex , LangChain ), and SQL.
  • Deep structural understanding of context window limitations ("lost in the middle" phenomenon), multi-modal token dynamics, network data transfer speeds, and cloud memory spaces.
  • Mandatory certification: Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud vendor (AWS/GCP/Azure).
Preferred Qualifications
  • Prior experience implementing Graph RAG frameworks utilizing native knowledge graphs (e.g., Neo4j) to map complex corporate data relationship networks.
  • Familiarity with fine-tuning open-source text embedding models specifically optimized for industry-specific terminology or legacy product schemas.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RAG/LLM Specialist_ Manager/Sr. Manager
RAG/LLM Specialist_ Manager/Sr. Manager

EXL • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Analytics Engineer-Senior AI / RAG Platform Engineer
Analytics Engineer-Senior AI / RAG Platform Engineer

Trigyn Technologies Limited. • Delhi

On-site
INR 2,000,000 - 3,500,000
Retrieval-Augmented Generation (RAG) Engineer
Retrieval-Augmented Generation (RAG) Engineer

Yantran • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior RAG & Knowledge Systems Engineer
Senior RAG & Knowledge Systems Engineer

The Enterprise • Hyderabad

On-site
INR 1,500,000 - 2,000,000
RAG AI Developer (LLM + Retrieval) – EdTech
RAG AI Developer (LLM + Retrieval) – EdTech

AP Guru • Mumbai

On-site
INR 1,100,000 - 1,700,000
Senior RAG Document AI Engineer
Senior RAG Document AI Engineer

3across • Hyderabad

On-site
INR 3,500,000 - 5,500,000
Senior RAG - Document AI Engineer
Senior RAG - Document AI Engineer

Chryselys • Hyderabad, Chennai District

On-site
INR 300,000 - 600,000
AI Engineer (Generative AI / RAG)
AI Engineer (Generative AI / RAG)

Xminds Infotech Pvt Ltd • Thiruvananthapuram

On-site
INR 1,500,000 - 2,100,000
AI Engineer – RAG & Agentic AI
AI Engineer – RAG & Agentic AI

IntraEdge • Hyderabad

Hybrid
INR 4,000,000 - 6,500,000
Lead AI/ML Engineer ( NLP, Transformers, Vector Databases, and RAG),
Lead AI/ML Engineer ( NLP, Transformers, Vector Databases, and RAG),

Optum India • Bengaluru

On-site
INR 4,000,000 - 7,000,000