← All projects

Reference Architecture

Agentic Enterprise RAG

Agentic RAG reference system using LangGraph orchestration, Weaviate hybrid retrieval, local/cloud LLM modes and a FastAPI backend.

GitHub ↗

Overview

Explores document retrieval and bounded tool-using answer workflows through a shared FastAPI backend.

Architecture

  • Document processing and embeddings feed Weaviate vector/BM25 hybrid retrieval (backend/app/core/vector_store.py).
  • A LangGraph decision/tool loop preserves retrieved source context for the final response (backend/app/agent/graph.py and nodes.py).
  • Provider configuration selects Ollama or OpenAI models (backend/app/core/llm_provider.py).

Engineering decisions

  • Combine semantic retrieval with keyword matching.
  • Separate LLM provider configuration from retrieval and orchestration.
  • Cap graph decisions at five steps and preserve source metadata across tool calls.

What is implemented

  • Document upload and query endpoints in backend/app/api/.
  • Weaviate hybrid retrieval and LangGraph tool orchestration.
  • Language detection and English/German prompts in backend/app/core/language_detect.py.
  • Optional regex query masking and estimated token-cost tracking.

Current limitations

  • Local mode includes cloud fallback when OpenAI credentials are configured; local-only privacy is not guaranteed.
  • Regex masking does not establish GDPR compliance or enterprise authorization.
  • Costs are estimates. Deployment configuration and evaluation utilities do not establish a verified live service or answer-quality benchmark.

Stack

  • Python
  • LangGraph
  • Weaviate
  • RAG
  • FastAPI