← All projects

Portfolio Prototype

Self-Improving LLM Knowledge Base

RAG portfolio system with FAISS/BM25 hybrid retrieval, persistent Q&A memory, generated summary notes and evaluation utilities.

GitHub ↗

Overview

Explores hybrid retrieval and persistent interaction records for a Markdown knowledge base.

Architecture

  • Markdown parsing and chunking feed FAISS and BM25 indexes (src/ingestion/ and src/retrieval/).
  • Weighted reciprocal-rank fusion combines retrieval results before context-based LLM generation (src/retrieval/hybrid.py and src/llm/reasoning.py).
  • Q&A records and generated summary notes are persisted to disk (src/memory/store.py).

Engineering decisions

  • Offer dense, sparse and hybrid retrieval for comparison.
  • Separate retrieval, generation and disk-backed interaction storage.
  • Allow optional reranking and MLflow evaluation logging.

What is implemented

  • Markdown ingestion, FAISS/BM25 retrieval and weighted RRF.
  • Optional cross-encoder reranking and context-based generation.
  • JSONL Q&A storage and Markdown summary-note writes.
  • Streamlit UI, Recall@K/MRR metrics, heuristic answer scoring and optional MLflow tracking.

Current limitations

  • Summary notes are written to disk but are not automatically re-indexed; they require an explicit ingestion cycle to enter retrieval.
  • The query path does not automatically use stored conversation history as generation input.
  • Heuristic answer scoring is not hallucination detection. The LLM-as-judge utility generates a prompt rather than a validated benchmark.
  • Optional MLflow logging does not establish fully reproducible experiments or continuous autonomous improvement.

Stack

  • Python
  • RAG
  • FAISS
  • BM25
  • RRF
  • Streamlit
  • MLflow