A Retrieval-Augmented Generation (RAG) app for querying your own PDFs/notes, running entirely on local compute except for the final answer generation call.
Pipeline: Sentence-Transformer embeddings → local FAISS vector search → cross-encoder reranking → grounded answer generation via the Gemini API → Streamlit UI with source-passage citations.
python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env
# then edit .env and paste your Gemini API key
# (get a free one at https://aistudio.google.com/apikey)streamlit run app.py- In the sidebar, upload one or more
.pdf,.txt, or.mdfiles. - Click Build / Rebuild Index — this chunks the text, embeds it with Sentence-Transformers, and builds a local FAISS index.
- Ask questions in the chat box. Each answer includes an expandable Sources section showing which passages (and page numbers) were used.
- Ingest (
src/ingest.py): documents are loaded, split into overlapping chunks with LangChain'sRecursiveCharacterTextSplitter, embedded with a Sentence-Transformer model (all-MiniLM-L6-v2by default), and stored in a local FAISSIndexFlatIPindex (cosine similarity via normalized vectors). - Retrieve (
src/retriever.py): the query is embedded and FAISS returns the top-N candidate chunks. A cross-encoder (cross-encoder/ms-marco-MiniLM-L-6-v2) then rescoring the (query, chunk) pairs directly, which is more accurate than embedding similarity alone, and the top-K are kept. - Generate (
src/generator.py): the reranked passages are inserted into a grounded prompt and sent to the Gemini API, which is instructed to answer only from the given context and cite sources inline. - UI (
app.py): Streamlit handles uploads, indexing, chat history, and rendering of the cited source passages.
Re-run indexing any time you add or change documents in data/raw_docs/ —
either via the sidebar button or:
python -m src.ingest- Everything except the final Gemini call runs locally — no cloud vector DB, no external embedding API.
- Swap
EMBEDDING_MODEL,CROSS_ENCODER_MODEL,GEMINI_MODEL, chunk sizes, and top-k values in.envwithout touching code.