Skip to content
AbdullahWali007Public

About

Context-aware multimodal RAG system with 4-agent orchestration. Dynamically filters vector search by user environment, reranks retrieved chunks with cross-encoder, and synthesizes contextually accurate answers. Includes Ragas evaluation and LangSmith observability. Built with FastAPI, Qdrant, LlamaIndex.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

DocuSync AI: Context-Aware Multimodal RAG Engine

Python FastAPI LlamaIndex Qdrant Gemini

DocuSync AI is an enterprise-grade, multi-agent Retrieval-Augmented Generation (RAG) system. Unlike standard RAG pipelines that blindly retrieve text, DocuSync utilizes a 4-Agent Orchestration Flow to read a user's environment profile (OS, Framework, Language), dynamically rewrite search queries, strictly filter vector databases, and synthesize highly contextual answers with zero hallucination.

Automated Evaluation Metrics (Ragas & LangSmith)

This system was mathematically evaluated using the ragas framework against a ground-truth dataset, proving its enterprise readiness and guardrail efficacy.

  • Faithfulness (Hallucination Guardrail): 97.06% (System successfully refuses to invent code when documentation is missing for a specific OS/Framework).
  • Context Precision: 100% (On Valid Contexts) (Cross-Encoder reranking ensures the top returned node is the exact match for the user's environment).
  • Observability: 100% of pipeline executions, token usage, and cross-encoder latencies are traced and monitored via LangSmith.

System Architecture

graph TD
    subgraph Frontend [UI Layer]
        UI[Web Interface]
        Profile[User Persona Profile]
    end

    subgraph API [FastAPI Backend]
        Router[API Router]
    end

    subgraph Agentic_Flow [4-Agent Orchestration]
        A1[Agent A: Context Extractor]
        A2[Agent B: Query Rewriter & Retriever]
        A3[Agent C: Cross-Encoder Reranker]
        A4[Agent D: Contextual Synthesizer]
    end

    subgraph Data_Layer [Vector Infrastructure]
        Parser[Context-Aware Chunking]
        Embed[BAAI/bge-large-en-v1.5]
        DB[(Qdrant Vector Database)]
    end

    %% Ingestion Flow
    Docs((Raw Docs)) --> Parser
    Parser -->|Intelligent Metadata Tagging| Embed
    Embed --> DB

    %% Query Flow
    UI -->|Raw Query| Router
    Profile -->|JSON Context| Router
    Router --> A1
    
    A1 -->|Structured Constraints| A2
    A2 -->|Optimized Query + Strict Filters| DB
    DB -->|Top 20 Nodes| A3
    A3 -->|Top 5 High-Precision Nodes| A4
    
    A4 -->|Gemini 2.5 Synthesis| Router
    Router -->|Final Answer| UI

    %% Styling
    classDef primary fill:#0f172a,stroke:#38bdf8,stroke-width:2px,color:#f8fafc;
    classDef secondary fill:#1e293b,stroke:#a78bfa,stroke-width:2px,color:#f8fafc;
    
    class A1,A2,A3,A4 primary;
    class DB,Embed,Parser secondary;

Loading

Core Features

  1. Intelligent Ingestion (document_parser.py): Deep content analysis that scans markdown headers and code blocks to dynamically assign os, language, and framework metadata tags to vector chunks before embedding.
  2. Dynamic Metadata Filtering: Prevents context contamination. A macOS/React developer will never be served Windows/Python installation instructions.
  3. Cross-Encoder Reranking: Utilizes HuggingFace ms-marco-MiniLM-L-6-v2 to mathematically score and rerank retrieved nodes, boosting precision by filtering out semantically similar but contextually irrelevant chunks.
  4. Asynchronous Backend: Built on FastAPI with asyncio.to_thread execution, ensuring the API remains highly responsive during computationally heavy ML embedding and reranking tasks.

Tech Stack

  • LLM: Google Gemini 2.5 Flash/ (You can also switch to ollama if Gemini API is not feasible)
  • Orchestration: LlamaIndex
  • Vector Database: Qdrant (Local Backend)
  • Embeddings: HuggingFace (BAAI/bge-large-en-v1.5)
  • Reranker: Sentence Transformers (cross-encoder/ms-marco-MiniLM-L-6-v2)
  • Backend: FastAPI, Uvicorn, Pydantic
  • MLOps / Evaluation: LangSmith, Ragas, Datasets
  • Frontend: Vanilla HTML/CSS/JS (Glassmorphism UI)

Quick Start

1. Prerequisites

Ensure you have Python 3.10+ installed. Clone the repository and install the dependencies.

git clone https://github.com/AbdullahWali007/DocuSync.git
cd DocuSync
pip install -r requirements.txt

2. Environment Variables

Create a .env file in the root directory:

GEMINI_API_KEY="your_gemini_key"
LANGCHAIN_API_KEY="your_langsmith_key"   # Only required for Phase 4 tracing
LANGCHAIN_PROJECT="DocuSync_AI_Production"

3. Initialize the Vector Database (Phase 1)

Run the ingestion script to parse your documents, generate embeddings, and populate Qdrant.

python ingest.py

Note: This single command internally imports and uses the supporting modules config.py, document_parser.py, and qdrant_setup.py – you do not need to run them manually.

4. Start the Application (Phase 5)

Spin up the FastAPI backend server:

python phase5_api.py

Once the server is running on http://localhost:8000, open index.html in your web browser to interact with the Multimodal UI.


Pipeline Execution Order (Detailed)

For those who wish to run the optional evaluation or tracing phases, below is the full recommended sequence:

Step File Purpose
1 ingest.py (Mandatory) Loads, chunks, enriches metadata, embeds, and stores documents in Qdrant.
2 (Optional) phase2_baseline.py Runs a set of benchmark queries against a simple Python-only retriever; exports results to phase2_control_state.json. Useful for before/after comparisons.
3 (Optional) phase3_orchestration.py Can be run directly (has __main__) to test a single query against the 4‑Agent pipeline without the API. Mainly used for development.
4 (Optional) phase4_tracing.py Wraps the orchestrator with LangSmith @traceable decorators and executes a sample query. Requires LANGCHAIN_API_KEY.
5 (Optional) phase4_ragas_eval.py Runs the Ragas evaluation suite against a ground‑truth dataset; outputs a CSV with per‑sample metrics (faithfulness, context precision, answer relevancy).
6 phase5_api.py (Production) Starts the FastAPI server. All other phases are integrated; this is the final service entrypoint.

Important: The supporting library files (config.py, document_parser.py, qdrant_setup.py) are not executed directly; they are imported by the scripts above.


Evaluation & Observability (Phases 4)

To run the automated evaluation and tracing:

# Trace a query with LangSmith (ensure LANGCHAIN_API_KEY is set)
python phase4_tracing.py

# Evaluate the pipeline with Ragas on a ground-truth dataset
python phase4_ragas_eval.py

Results are saved locally (phase4_ragas_metrics.csv) and visible in your LangSmith dashboard.


Engineered by M.Abdullah Wali.

About

Context-aware multimodal RAG system with 4-agent orchestration. Dynamically filters vector search by user environment, reranks retrieved chunks with cross-encoder, and synthesizes contextually accurate answers. Includes Ragas evaluation and LangSmith observability. Built with FastAPI, Qdrant, LlamaIndex.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages