Mentat Commons

Designing decisions for collective intelligence

Back to all services

RAG Architecture

Semantic chunking, metadata injection, and vector search retrievers.

Published: 2026-08-01•Last Updated: 2026-08-07

What is this service?

Construct high-precision Retrieval-Augmented Generation systems using semantic chunking and advanced hybrid search.

Who is this for?

Product companies and enterprise teams with large unstructured document corpuses needing accurate question answering.

Pricing & Engagement

  • Starting from: ₹18,00,000
  • Typical range: ₹18,00,000 - ₹40,00,000
  • Pricing Model: Phase-based implementation
  • Engagement timeline: 4 to 8 weeks

Typical Problems We Resolve

!

Naïve text chunking splitting critical paragraphs and losing context.

!

Irrelevant documents retrieved by vector search polluting the LLM context window.

!

Out-of-date vector databases missing latest document updates.

Key Deliverables

Custom Document Parsing & Semantic Chunking Pipeline
Hybrid Search Retrieval Engine (BM25 + Dense Embeddings)
Automated Index Ingestion & Syncer Daemon

Technology Stack

PythonQdrantpgvectorCohere RerankLlamaIndexDocker

Expected Outcomes

✓

Retrieval accuracy (MRR) improved from 60% to 92%.

✓

LLM context token waste reduced by 40%.

✓

Near real-time vector database updates.

Next Steps & Engagement

Set up a deep-dive call to inspect sample document structures and define parsing constraints.