Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Β | Β | |||
Β | Β | |||
Β | Β | |||
Β | Β | |||
Repository files navigation
# π§ PDF Chatbot using Sentence Transformers and FAISS This project builds a simple AI-powered chatbot that allows users to **ask questions about any PDF document**. It uses **semantic similarity search** powered by **Sentence Transformers** and **FAISS** to retrieve relevant sections from the PDF. --- ## π Features - π Extracts text from PDF using PyMuPDF (`fitz`) - π§© Splits long PDF text into manageable chunks - π Generates semantic embeddings using [`all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) - β‘ Stores and retrieves embeddings using [FAISS](https://github.com/facebookresearch/faiss) for fast vector search - β Supports natural language queries to search the PDF content --- ## π Project Structure --- ## π οΈ Requirements - Python 3.8+ - `sentence-transformers` - `faiss-cpu` - `PyMuPDF` (`fitz`) Install using: ```bash pip install -r requirements.txt How It Works Text Extraction: Reads the entire PDF and extracts text using PyMuPDF. Chunking: Splits the text into equal-sized chunks (default: 512 words). Embedding Generation: Converts each chunk into a dense vector using a SentenceTransformer model. Vector Indexing: Uses FAISS to store and perform similarity search on embeddings. Search: Given a query, the system retrieves the most semantically similar chunks from the PDF.