RAG Architect

Zoho

United States

Remote

USD 150,000 - 210,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Fyerx is seeking an experienced Context Window Optimization / RAG Architect to own enterprise generative AI retrieval performance, accuracy, and cost metrics. You will design high-throughput retrieval systems, optimize semantic parsing, build re-ranking pipelines, and engineer caching grids for data-grounded AI responses with minimal latency and maximum token efficiency.

As the lead, you will architect end-to-end RAG pipelines, optimize chunking and context usage, implement semantic caching, and

Qualifications

  • 6 to 10 years of enterprise data engineering, database design, or search engine engineering experience.

Responsibilities

  • Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, with document parsing and multi-vector lookups.
  • Optimize context window utilization with smart chunking, sentence-window retrievals, and sliding window strategies.
  • Build high-performance re-ranking layers using cross-encoders (e.g., Cohere Rerank, BGE-Reranker).
  • Implement automated semantic caching architectures to reduce API token costs and latency.
  • Establish automated data chunking pipelines to parse PDFs, wikis, and SQL outputs into vector targets.
  • Govern vector similarity spaces by fine-tuning hybrid search combining dense embeddings with BM25.
  • Audit context-level hallucination rates and logs to improve precision and recall.

Skills

Python
Embedding models
Context window optimization
Performance engineering

Education

Cloud Data/Database Engineer Certification
Specialty Analytics credential (AWS/GCP/Azure)

Tools

Pinecone
Milvus
Weaviate
LlamaIndex
LangChain
BM25
GPTCache

Job description

  • Total Experience Required: 6 to 10 years
  • Relevant Experience Required: 3+ years of dedicated experience designing production-grade Retrieval-Augmented Generation (RAG) architectures and optimizing LLM token throughput
Job Summary

We are seeking an experienced Context Window Optimization / RAG Architect to take full ownership of our enterprise generative AI retrieval performance, accuracy, and operational cost metrics. The ideal candidate will design high-throughput knowledge retrieval systems, optimize semantic context parsing, build custom re-ranking pipelines, and engineer caching grids to deliver data-grounded AI responses with minimal latency and maximum token efficiency.

Key Responsibilities

  • Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, building structures for document parsing, semantic metadata enrichment, and multi-vector lookups.
  • Optimize context window utilization patterns, designing smart parent-child chunking models, sentence-window retrievals, and sliding window strategies to eliminate irrelevant text tokens.
  • Build high-performance re-ranking layers, deploying machine learning cross-encoders (e.g., Cohere Rerank, BGE-Reranker) to score retrieved documents before feeding them into the LLM context pool.
  • Implement automated semantic caching architectures, utilizing caching layers (e.g., GPTCache) to capture recurring semantic queries, reducing API token expenditures and response latencies.
  • Establish automated data chunking pipelines, configuring ingestion routines to cleanly parse semi-structured and unstructured formats (PDFs, corporate wikis, SQL outputs) into clean vector targets.
  • Govern vector similarity spaces, fine-tuning hybrid search algorithms that cleanly combine dense semantic embeddings with sparse keyword token indexes (BM25).
  • Audit context-level hallucination rates and accuracy logs, tracking precision metrics, retrieval recall bounds, and processing speeds to systematically eliminate incorrect model generations.
Requirements
  • 6 to 10 years of enterprise data engineering, database design, or search engine engineering experience, with 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
  • Strong technical mastery of Python, vector databases (Pinecone, Milvus , Weaviate ), text embedding models, open-source orchestration tools ( LlamaIndex , LangChain ), and SQL.
  • Deep structural understanding of context window limitations ("lost in the middle" phenomenon), multi-modal token dynamics, network data transfer speeds, and cloud memory spaces.
  • Mandatory certification: Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud vendor (AWS/GCP/Azure).
Preferred Qualifications
  • Prior experience implementing Graph RAG frameworks utilizing native knowledge graphs (e.g., Neo4j) to map complex corporate data relationship networks.
  • Familiarity with fine-tuning open-source text embedding models specifically optimized for industry-specific terminology or legacy product schemas.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Retrieval-Augmented Generation (RAG) System - Senior Software Developer
Retrieval-Augmented Generation (RAG) System - Senior Software Developer

Elsevier • Philadelphia

On-site
USD 120,000 - 150,000
RAG Architect: High-Throughput Knowledge Retrieval
RAG Architect: High-Throughput Knowledge Retrieval

Zoho • United States

Remote
USD 150,000 - 210,000
ML Engineer II
ML Engineer II

Compunnel, Inc. • Mason (OH)

On-site
USD 100,000 - 130,000
Senior GenAI Engineer
Senior GenAI Engineer

Global Business Ser. 4u • New York (NY)

On-site
USD 150,000 - 220,000
Senior Polyglot Engineer
Senior Polyglot Engineer

Fanisko • United States

On-site
USD 120,000 - 160,000
Gen. AI Engineer
Gen. AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 140,000 - 190,000
Senior RAG Architect for Enterprise AI
Senior RAG Architect for Enterprise AI

MCI • United States

On-site
USD 150,000 - 230,000
Senior AI Engineer – Generative AI / RAG
Senior AI Engineer – Generative AI / RAG

Ampcus Inc • Chantilly (VA)

On-site
USD 140,000 - 200,000
Programmer
Programmer

Atlas • Village of Tarrytown (NY)

On-site
USD 180,000 - 260,000
Ai Software Architect Llm Agentic Systems Msys Tech India Pvt Ltd Chennai
Ai Software Architect Llm Agentic Systems Msys Tech India Pvt Ltd Chennai

Vibehackers • United States

On-site
USD 180,000 - 260,000