Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

# 🧠 PDF Chatbot using Sentence Transformers and FAISS

This project builds a simple AI-powered chatbot that allows users to **ask questions about any PDF document**. It uses **semantic similarity search** powered by **Sentence Transformers** and **FAISS** to retrieve relevant sections from the PDF.

---

## πŸš€ Features

- πŸ“„ Extracts text from PDF using PyMuPDF (`fitz`)
- 🧩 Splits long PDF text into manageable chunks
- πŸ” Generates semantic embeddings using [`all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2)
- ⚑ Stores and retrieves embeddings using [FAISS](https://github.com/facebookresearch/faiss) for fast vector search
- ❓ Supports natural language queries to search the PDF content

---

## πŸ“ Project Structure


---

## πŸ› οΈ Requirements

- Python 3.8+
- `sentence-transformers`
- `faiss-cpu`
- `PyMuPDF` (`fitz`)

Install using:

```bash
pip install -r requirements.txt
 How It Works
Text Extraction: Reads the entire PDF and extracts text using PyMuPDF.

Chunking: Splits the text into equal-sized chunks (default: 512 words).

Embedding Generation: Converts each chunk into a dense vector using a SentenceTransformer model.

Vector Indexing: Uses FAISS to store and perform similarity search on embeddings.

Search: Given a query, the system retrieves the most semantically similar chunks from the PDF.

About

No description or website provided.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages