Reference Architecture
Agentic Enterprise RAG
Agentic RAG reference system using LangGraph orchestration, Weaviate hybrid retrieval, local/cloud LLM modes and a FastAPI backend.
GitHub ↗Overview
Explores document retrieval and bounded tool-using answer workflows through a shared FastAPI backend.
Architecture
- Document processing and embeddings feed Weaviate vector/BM25 hybrid retrieval (backend/app/core/vector_store.py).
- A LangGraph decision/tool loop preserves retrieved source context for the final response (backend/app/agent/graph.py and nodes.py).
- Provider configuration selects Ollama or OpenAI models (backend/app/core/llm_provider.py).
Engineering decisions
- Combine semantic retrieval with keyword matching.
- Separate LLM provider configuration from retrieval and orchestration.
- Cap graph decisions at five steps and preserve source metadata across tool calls.
What is implemented
- Document upload and query endpoints in backend/app/api/.
- Weaviate hybrid retrieval and LangGraph tool orchestration.
- Language detection and English/German prompts in backend/app/core/language_detect.py.
- Optional regex query masking and estimated token-cost tracking.
Current limitations
- Local mode includes cloud fallback when OpenAI credentials are configured; local-only privacy is not guaranteed.
- Regex masking does not establish GDPR compliance or enterprise authorization.
- Costs are estimates. Deployment configuration and evaluation utilities do not establish a verified live service or answer-quality benchmark.