DocuSync AI is an enterprise-grade, multi-agent Retrieval-Augmented Generation (RAG) system. Unlike standard RAG pipelines that blindly retrieve text, DocuSync utilizes a 4-Agent Orchestration Flow to read a user's environment profile (OS, Framework, Language), dynamically rewrite search queries, strictly filter vector databases, and synthesize highly contextual answers with zero hallucination.
This system was mathematically evaluated using the ragas framework against a ground-truth dataset, proving its enterprise readiness and guardrail efficacy.
- Faithfulness (Hallucination Guardrail): 97.06% (System successfully refuses to invent code when documentation is missing for a specific OS/Framework).
- Context Precision: 100% (On Valid Contexts) (Cross-Encoder reranking ensures the top returned node is the exact match for the user's environment).
- Observability: 100% of pipeline executions, token usage, and cross-encoder latencies are traced and monitored via LangSmith.
graph TD
subgraph Frontend [UI Layer]
UI[Web Interface]
Profile[User Persona Profile]
end
subgraph API [FastAPI Backend]
Router[API Router]
end
subgraph Agentic_Flow [4-Agent Orchestration]
A1[Agent A: Context Extractor]
A2[Agent B: Query Rewriter & Retriever]
A3[Agent C: Cross-Encoder Reranker]
A4[Agent D: Contextual Synthesizer]
end
subgraph Data_Layer [Vector Infrastructure]
Parser[Context-Aware Chunking]
Embed[BAAI/bge-large-en-v1.5]
DB[(Qdrant Vector Database)]
end
%% Ingestion Flow
Docs((Raw Docs)) --> Parser
Parser -->|Intelligent Metadata Tagging| Embed
Embed --> DB
%% Query Flow
UI -->|Raw Query| Router
Profile -->|JSON Context| Router
Router --> A1
A1 -->|Structured Constraints| A2
A2 -->|Optimized Query + Strict Filters| DB
DB -->|Top 20 Nodes| A3
A3 -->|Top 5 High-Precision Nodes| A4
A4 -->|Gemini 2.5 Synthesis| Router
Router -->|Final Answer| UI
%% Styling
classDef primary fill:#0f172a,stroke:#38bdf8,stroke-width:2px,color:#f8fafc;
classDef secondary fill:#1e293b,stroke:#a78bfa,stroke-width:2px,color:#f8fafc;
class A1,A2,A3,A4 primary;
class DB,Embed,Parser secondary;
- Intelligent Ingestion (
document_parser.py): Deep content analysis that scans markdown headers and code blocks to dynamically assignos,language, andframeworkmetadata tags to vector chunks before embedding. - Dynamic Metadata Filtering: Prevents context contamination. A macOS/React developer will never be served Windows/Python installation instructions.
- Cross-Encoder Reranking: Utilizes HuggingFace
ms-marco-MiniLM-L-6-v2to mathematically score and rerank retrieved nodes, boosting precision by filtering out semantically similar but contextually irrelevant chunks. - Asynchronous Backend: Built on FastAPI with
asyncio.to_threadexecution, ensuring the API remains highly responsive during computationally heavy ML embedding and reranking tasks.
- LLM: Google Gemini 2.5 Flash/ (You can also switch to ollama if Gemini API is not feasible)
- Orchestration: LlamaIndex
- Vector Database: Qdrant (Local Backend)
- Embeddings: HuggingFace (
BAAI/bge-large-en-v1.5) - Reranker: Sentence Transformers (
cross-encoder/ms-marco-MiniLM-L-6-v2) - Backend: FastAPI, Uvicorn, Pydantic
- MLOps / Evaluation: LangSmith, Ragas, Datasets
- Frontend: Vanilla HTML/CSS/JS (Glassmorphism UI)
Ensure you have Python 3.10+ installed. Clone the repository and install the dependencies.
git clone https://github.com/AbdullahWali007/DocuSync.git
cd DocuSync
pip install -r requirements.txt
Create a .env file in the root directory:
GEMINI_API_KEY="your_gemini_key"
LANGCHAIN_API_KEY="your_langsmith_key" # Only required for Phase 4 tracing
LANGCHAIN_PROJECT="DocuSync_AI_Production"
Run the ingestion script to parse your documents, generate embeddings, and populate Qdrant.
python ingest.py
Note: This single command internally imports and uses the supporting modules config.py, document_parser.py, and qdrant_setup.py – you do not need to run them manually.
Spin up the FastAPI backend server:
python phase5_api.py
Once the server is running on http://localhost:8000, open index.html in your web browser to interact with the Multimodal UI.
For those who wish to run the optional evaluation or tracing phases, below is the full recommended sequence:
| Step | File | Purpose |
|---|---|---|
| 1 | ingest.py |
(Mandatory) Loads, chunks, enriches metadata, embeds, and stores documents in Qdrant. |
| 2 (Optional) | phase2_baseline.py |
Runs a set of benchmark queries against a simple Python-only retriever; exports results to phase2_control_state.json. Useful for before/after comparisons. |
| 3 (Optional) | phase3_orchestration.py |
Can be run directly (has __main__) to test a single query against the 4‑Agent pipeline without the API. Mainly used for development. |
| 4 (Optional) | phase4_tracing.py |
Wraps the orchestrator with LangSmith @traceable decorators and executes a sample query. Requires LANGCHAIN_API_KEY. |
| 5 (Optional) | phase4_ragas_eval.py |
Runs the Ragas evaluation suite against a ground‑truth dataset; outputs a CSV with per‑sample metrics (faithfulness, context precision, answer relevancy). |
| 6 | phase5_api.py |
(Production) Starts the FastAPI server. All other phases are integrated; this is the final service entrypoint. |
Important: The supporting library files (
config.py,document_parser.py,qdrant_setup.py) are not executed directly; they are imported by the scripts above.
To run the automated evaluation and tracing:
# Trace a query with LangSmith (ensure LANGCHAIN_API_KEY is set)
python phase4_tracing.py
# Evaluate the pipeline with Ragas on a ground-truth dataset
python phase4_ragas_eval.py
Results are saved locally (phase4_ragas_metrics.csv) and visible in your LangSmith dashboard.
Engineered by M.Abdullah Wali.